Binary Translator

Text to binary and back in UTF-8, UTF-16, code points or Latin-1, with all four shown so it is clear the answer depends on the encoding.

Enable JavaScript to customise; default output below.

Direction

Decoding accepts any separator: spaces, commas, newlines. A 0b, 0x or 0o prefix is ignored.

Encoding

UTF-8 is what a file holds. UTF-16 is what a JavaScript string is made of, which is what charCodeAt gives you. A code point is the number behind U+XXXX.

Base
Live preview binary.txt
01001000 11000011 10101001 00100000 11110000 10011111 10011000 10000000

Characters in                  4
Units out                      8, UTF-8 bytes, which is what a file or a request holds
Bits                           64

The same text, the other ways
  utf-8                        8 units  ← yours
48 C3 A9 20 F0 9F 98 80
  utf-16                       5 units
48 E9 20 D83D DE00
  code-point                   4 units
48 E9 20 01F600
  latin-1                      cannot: "😀" is U+1F600, past the 256 characters Latin-1 can hold.

"The binary of a character" depends on the encoding, and most
translators pick one silently. `charCodeAt` gives UTF-16 code units, so
an emoji comes out as two numbers that are halves of a surrogate pair
rather than as a character, and an accented letter comes out as one
sixteen-bit value where UTF-8 would have produced two bytes.

UTF-8 is what a file, a request and a database column hold. It uses one
byte for ASCII, two for most European letters, three for most other
scripts and four for emoji, which is why byte counts and character
counts differ and why a `varchar(255)` does not hold 255 characters of
Bengali.

UTF-16 is what a JavaScript string is made of in memory. A character
above U+FFFF takes two of its units, which is where lone surrogates,
broken string reversals and `"😀".length === 2` all come from.

A code point is the number Unicode assigns: U+1F600 is 128,512. Neither
UTF-8 nor UTF-16 stores it directly at that size, which is the
distinction between a character and its representation.

Latin-1 holds 256 characters and nothing else. It is worth having here
because a great deal of old data is in it, and because a file misread as
Latin-1 when it is UTF-8 is how "café" becomes "café": the two UTF-8
bytes get read as two separate characters.

The padding is presentation, not data. A byte is eight binary digits
because a byte is eight bits, and a UTF-16 unit is sixteen. Unpadded
output is ambiguous: "1 10" could be two units or one, which is why the
decoder here needs separators.

This is not encryption. Binary is a way of writing a number down, and
anything written in it can be read back by anybody. If something needs
to be hidden, this and base64 are both the wrong tool.

Output is valid and updates as you type.

“The binary of a character” is not a question with one answer. It depends on the encoding, and most translators pick one without telling you which.

The usual choice is charCodeAt, which gives UTF-16 code units. Under that, an emoji arrives as two numbers that are halves of a surrogate pair rather than a character, and é comes out as one sixteen-bit value where a UTF-8 encoder would have produced two bytes. Neither is wrong; they are answers to different questions.

So all four are shown here. UTF-8 is what a file, a request and a database column hold. UTF-16 is what a JavaScript string is made of. A code point is the number Unicode assigns, which is what U+1F600 means. Latin-1 is where a great deal of old data still lives.

How to use

  1. Type text, or paste numbers to decode.
  2. Pick the encoding. UTF-8 unless you know you want another.
  3. Pick the base. Decoding accepts any separator and ignores a 0b, 0x or 0o prefix.

Example

Hé 😀 in UTF-8 binary:

01001000 11000011 10101001 00100000 11110000 10011111 10011000 10000000

Characters in                  4
Units out                      8, UTF-8 bytes, which is what a file or a request holds
Bits                           64

The same text, the other ways
  utf-8                        8 units  ← yours
48 C3 A9 20 F0 9F 98 80
  utf-16                       5 units
48 E9 20 D83D DE00
  code-point                   4 units
48 E9 20 01F600
  latin-1                      cannot: "😀" is U+1F600, past the 256 characters Latin-1 can hold.

Four characters. Eight bytes, five code units, four code points, and impossible in Latin-1. Every one of those numbers is correct about something different.

Pitfalls

charCodeAt is not the binary of a character. It returns a UTF-16 code unit, so anything above U+FFFF comes back as half of a surrogate pair. Use TextEncoder for bytes and codePointAt for code points.

A byte count is not a character count. UTF-8 uses one byte for ASCII, two for most European letters, three for most other scripts and four for emoji. That is why a varchar(255) does not hold 255 characters of Bengali, and why truncating at a byte boundary can cut a character in half.

Unpadded binary is ambiguous. 1 10 could be two units or one. The padding here is eight digits for a byte and sixteen for a UTF-16 unit, because that is how many bits each holds, and the decoder needs separators for the same reason.

Latin-1 misreadings are recognisable. A UTF-8 file read as Latin-1 turns café into café: the two bytes for é get read as two separate characters. If you see that pattern, the data is fine and something decoded it wrongly.

MySQL’s utf8 is not UTF-8. It holds at most three bytes a character, so it cannot store emoji. The one you want is utf8mb4, and a column in the older one silently truncates or errors on a four-byte character.

This is not encryption. Binary is a way of writing a number down. Anything written in it can be read back by anyone, and the same is true of base64. If something needs to be hidden, neither is the tool.

Octal is a trap in source code. A leading zero makes a number octal in several languages, so 010 is 8. It is included here because it appears in file permissions and escape sequences, not because it is a good way to store text.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored.

UTF-8 uses TextEncoder and TextDecoder, which are in every browser since 2017 and decode in fatal mode, so invalid byte sequences are refused rather than quietly replaced with question marks. UTF-16 uses charCodeAt, code points use codePointAt, and Latin-1 refuses anything past U+00FF and names the character it refused.

The round trip is checked in the test suite in all four bases for text containing an accent, a space and an emoji, which covers one, two and four byte sequences.

In PHP, mb_convert_encoding handles the conversions and unpack( 'C*', $string ) gives the bytes. str_split splits by byte, so on UTF-8 text it will cut multibyte characters apart; mb_str_split is the one that does not.

Frequently asked questions

Why is my emoji two numbers in one tool and one in another?
Because one is giving UTF-16 code units and the other code points. "😀".length is 2 in JavaScript for exactly this reason, and neither tool is broken.
Which encoding should I use for a database?
utf8mb4 in MySQL, UTF8 in Postgres. Anything else eventually meets a character it cannot store, and the failure usually appears in production with a customer’s name in it.
How many bytes is a character?
In UTF-8: one for ASCII, two up to U+07FF, three up to U+FFFF, four above that. So between one and four, which is why byte-based length limits and character-based ones disagree.
Can I decode binary I found somewhere?
Try UTF-8 first. If groups of eight decode to readable ASCII you are done; if you see accented nonsense, try Latin-1; if the groups are sixteen digits long, it is UTF-16.
What about base64?
Base64 encodes bytes as text rather than writing a number in another base, and it has its own alphabet and padding rules. The base64 encoder on this site is the tool for it.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.