Binary Translator
Text to binary and back in UTF-8, UTF-16, code points or Latin-1, with all four shown so it is clear the answer depends on the encoding.
01001000 11000011 10101001 00100000 11110000 10011111 10011000 10000000
Characters in 4
Units out 8, UTF-8 bytes, which is what a file or a request holds
Bits 64
The same text, the other ways
utf-8 8 units ← yours
48 C3 A9 20 F0 9F 98 80
utf-16 5 units
48 E9 20 D83D DE00
code-point 4 units
48 E9 20 01F600
latin-1 cannot: "😀" is U+1F600, past the 256 characters Latin-1 can hold.
"The binary of a character" depends on the encoding, and most
translators pick one silently. `charCodeAt` gives UTF-16 code units, so
an emoji comes out as two numbers that are halves of a surrogate pair
rather than as a character, and an accented letter comes out as one
sixteen-bit value where UTF-8 would have produced two bytes.
UTF-8 is what a file, a request and a database column hold. It uses one
byte for ASCII, two for most European letters, three for most other
scripts and four for emoji, which is why byte counts and character
counts differ and why a `varchar(255)` does not hold 255 characters of
Bengali.
UTF-16 is what a JavaScript string is made of in memory. A character
above U+FFFF takes two of its units, which is where lone surrogates,
broken string reversals and `"😀".length === 2` all come from.
A code point is the number Unicode assigns: U+1F600 is 128,512. Neither
UTF-8 nor UTF-16 stores it directly at that size, which is the
distinction between a character and its representation.
Latin-1 holds 256 characters and nothing else. It is worth having here
because a great deal of old data is in it, and because a file misread as
Latin-1 when it is UTF-8 is how "café" becomes "café": the two UTF-8
bytes get read as two separate characters.
The padding is presentation, not data. A byte is eight binary digits
because a byte is eight bits, and a UTF-16 unit is sixteen. Unpadded
output is ambiguous: "1 10" could be two units or one, which is why the
decoder here needs separators.
This is not encryption. Binary is a way of writing a number down, and
anything written in it can be read back by anybody. If something needs
to be hidden, this and base64 are both the wrong tool.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
“The binary of a character” is not a question with one answer. It depends on the encoding, and most translators pick one without telling you which.
The usual choice is charCodeAt, which gives UTF-16 code units. Under that, an emoji arrives as two
numbers that are halves of a surrogate pair rather than a character, and é comes out as one
sixteen-bit value where a UTF-8 encoder would have produced two bytes. Neither is wrong; they are
answers to different questions.
So all four are shown here. UTF-8 is what a file, a request and a database column hold. UTF-16 is
what a JavaScript string is made of. A code point is the number Unicode assigns, which is what
U+1F600 means. Latin-1 is where a great deal of old data still lives.
How to use
- Type text, or paste numbers to decode.
- Pick the encoding. UTF-8 unless you know you want another.
- Pick the base. Decoding accepts any separator and ignores a
0b,0xor0oprefix.
Example
Hé 😀 in UTF-8 binary:
01001000 11000011 10101001 00100000 11110000 10011111 10011000 10000000
Characters in 4
Units out 8, UTF-8 bytes, which is what a file or a request holds
Bits 64
The same text, the other ways
utf-8 8 units ← yours
48 C3 A9 20 F0 9F 98 80
utf-16 5 units
48 E9 20 D83D DE00
code-point 4 units
48 E9 20 01F600
latin-1 cannot: "😀" is U+1F600, past the 256 characters Latin-1 can hold.
Four characters. Eight bytes, five code units, four code points, and impossible in Latin-1. Every one of those numbers is correct about something different.
Pitfalls
charCodeAt is not the binary of a character. It returns a UTF-16 code unit, so anything above
U+FFFF comes back as half of a surrogate pair. Use TextEncoder for bytes and codePointAt for
code points.
A byte count is not a character count. UTF-8 uses one byte for ASCII, two for most European
letters, three for most other scripts and four for emoji. That is why a varchar(255) does not hold
255 characters of Bengali, and why truncating at a byte boundary can cut a character in half.
Unpadded binary is ambiguous. 1 10 could be two units or one. The padding here is eight digits
for a byte and sixteen for a UTF-16 unit, because that is how many bits each holds, and the decoder
needs separators for the same reason.
Latin-1 misreadings are recognisable. A UTF-8 file read as Latin-1 turns café into café:
the two bytes for é get read as two separate characters. If you see that pattern, the data is fine
and something decoded it wrongly.
MySQL’s utf8 is not UTF-8. It holds at most three bytes a character, so it cannot store emoji.
The one you want is utf8mb4, and a column in the older one silently truncates or errors on a
four-byte character.
This is not encryption. Binary is a way of writing a number down. Anything written in it can be read back by anyone, and the same is true of base64. If something needs to be hidden, neither is the tool.
Octal is a trap in source code. A leading zero makes a number octal in several languages, so
010 is 8. It is included here because it appears in file permissions and escape sequences, not
because it is a good way to store text.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored.
UTF-8 uses TextEncoder and TextDecoder, which are in every browser since 2017 and decode in
fatal mode, so invalid byte sequences are refused rather than quietly replaced with question marks.
UTF-16 uses charCodeAt, code points use codePointAt, and Latin-1 refuses anything past U+00FF
and names the character it refused.
The round trip is checked in the test suite in all four bases for text containing an accent, a space and an emoji, which covers one, two and four byte sequences.
In PHP, mb_convert_encoding handles the conversions and unpack( 'C*', $string ) gives the bytes.
str_split splits by byte, so on UTF-8 text it will cut multibyte characters apart; mb_str_split
is the one that does not.
Frequently asked questions
Why is my emoji two numbers in one tool and one in another?
"😀".length is 2 in JavaScript
for exactly this reason, and neither tool is broken.Which encoding should I use for a database?
utf8mb4 in MySQL, UTF8 in Postgres. Anything else eventually meets a character it cannot store,
and the failure usually appears in production with a customer’s name in it.