Skip to content

Text to binary converter

Convert text to binary, hex, octal or decimal and back — with the character encoding made explicit, because that is what usually goes wrong.

This is the setting that matters. é is one byte in Latin-1 and two in UTF-8, so the same text gives different numbers and a mismatch is why decoded text comes back as gibberish.
Runs in your browser — nothing is sent anywhere.

The encoding is the whole point

Text is not bytes. Turning one into the other requires a choice of encoding, and that choice is the reason converted text comes back as gibberish.

Take é. In Latin-1 it is one byte, 233. In UTF-8 it is two, 195 and 169. Encode with one and decode with the other and you get é — the mojibake everyone has seen and few can explain. It is not corruption; it is two bytes being read as two characters instead of one.

So the encoding is a visible setting here rather than a hidden assumption. UTF-8 unless you know otherwise: it is what the web, JSON, modern databases and every current operating system use.

Padding is what makes the result decodable

The letter A is 65, which is 1000001 in binary — seven digits. B is 66, or 1000010. Run them together unpadded and you get 14 digits with no way to know they are two bytes rather than one long number.

Padding every byte to the same width — 8 digits in binary, 2 in hex, 3 in octal — is what allows a decoder to split the stream back up. Without either padding or separators the result is one-way. Decoding here will tell you so rather than guessing: if the digit count is not a multiple of the byte width, there is genuinely no correct answer.

Characters and bytes are different counts

This is the practical reason to reach for this tool. A database column declared as 20 characters is 20 bytes in many systems, and a name in Hindi, Arabic or Chinese uses three bytes per character in UTF-8. Twenty characters of Devanagari is sixty bytes, and the insert fails or truncates mid-character.

Both counts are shown, and the difference is called out when there is one.

Why hexadecimal is usually the better choice

Binary is the one people ask for and hex is the one they want. Two digits per byte instead of eight makes it four times shorter, it is what every debugger, hex editor, network capture and error message uses, and pairs of digits are easy to read off. Binary is worth it when you need to see individual bits — a flags field, a bitmask, a permission byte.

ASCII refuses rather than mangles

Choose ASCII and any character above 127 is refused with its code point, instead of being silently replaced with a question mark. Silent replacement is how a name loses its accent permanently, with nothing in the logs to say when.

01

Common questions

Why did my text come back as gibberish?

The encoding used to decode was not the one used to encode. é is one byte in Latin-1 and two in UTF-8, so reading UTF-8 as Latin-1 gives é. Try UTF-8 first.

Should I use binary or hex?

Hex, almost always — four times shorter, and it is what debuggers, hex editors and network captures use. Binary is for when you need to see individual bits.

Why does it need padding?

So a decoder can tell where each byte ends. A is 1000001, seven digits; without padding to eight, two bytes run together into an ambiguous string.

My text has more bytes than characters.

That is correct for UTF-8: accented, Cyrillic, Arabic, Devanagari and CJK characters take two to four bytes each. It is the byte count that matters for a length limit.

What does the ASCII option do differently?

It refuses anything above code point 127 and tells you which character, rather than silently substituting a question mark and losing it for good.

Can I decode without spaces between the bytes?

Yes, as long as every byte is padded to the same width. If the total digit count is not a multiple of that width you are told, because there is no correct way to split it.