Emoji Picker
Search 1,432 emoji by their official Unicode names, and see what any one is made of: code points, joiners, UTF-8 bytes and every escape.
This tool needs JavaScript: the list and the inspection both run in your browser.
rocket
Every emoji picker copies a character to the clipboard. This one also tells you what you just copied,
because that is where the bugs are: an emoji is often several code points rather than one, joined by an
invisible character, and JavaScript counts sixteen-bit units rather than characters, so "π¨βπ©βπ§".length
is 8.
Search by the official Unicode name, click to copy, and read what the thing is made of.
How to use
- Search by name, or filter by group. The names are Unicode’s own, so what your operating system calls an emoji is what finds it here.
- Click one to copy it and to inspect it.
- Or paste anything at all into the second box: the inspector works on emoji that are not in the list, on plain text, and on whatever your colleague pasted into that support ticket.
Example
The family emoji, taken apart:
π¨βπ©βπ§
Characters as people see them 1
Code points 5
JavaScript .length 8 UTF-16 units
UTF-8 bytes 18
Joined with ZWJ yes, 2 joiners
U+1F468 base π¨
U+200D joiner Zero width joiner. Glues the emoji either side of it into one glyph.
U+1F469 base π©
U+200D joiner Zero width joiner.
U+1F467 base π§
HTML 👨‍👩‍👧
JavaScript \u{1F468}\u{200D}\u{1F469}\u{200D}\u{1F467}
URL %F0%9F%91%A8%E2%80%8D%F0%9F%91%A9%E2%80%8D%F0%9F%91%A7
One thing on screen. Five code points. Eighteen bytes. Any validator that says “maximum 20 characters” has to pick one of those, and it usually picks the wrong one.
Pitfalls
utf8 in MySQL is not UTF-8. MySQL’s utf8 is three bytes a character and cannot store an emoji at
all: inserting one truncates the string at that point or throws, depending on the mode. The column has to
be utf8mb4, and so does the connection charset. WordPress has handled this since 4.2, which is why
wp-config.php sets DB_CHARSET to utf8mb4.
Counting characters is three different questions. Graphemes are what a person sees, code points are
what Unicode counts, and .length in JavaScript is UTF-16 units. A tweet-style limit wants graphemes; a
database column wants bytes; almost nothing wants .length, which is what almost everything uses.
Truncating by index cuts emoji in half. slice( 0, 20 ) on a string of emoji can end mid-sequence and
leave a lone surrogate, which renders as a replacement character and breaks JSON validity in some
parsers. Intl.Segmenter gives you graphemes to cut on instead.
The same emoji looks different everywhere. The code point is standard; the picture is the font. Apple, Google, Microsoft and Samsung all draw their own, and some of them have historically disagreed enough to change the meaning of a message.
Skin tone is a separate code point. A thumbs up with a tone is two code points, and stripping the modifier changes the character but not the base. The tools that “remove emoji” with a regular expression usually leave the modifiers behind.
Zero width joiner sequences degrade in public. A system without the glyph for a joined sequence shows its parts instead: a family becomes a man, a woman and a girl side by side. Nothing is broken; the font just does not have that one.
Emoji in a URL or a filename usually survive and occasionally do not. Percent-encoded they are fine; in a filename they depend on the filesystem’s normalisation, and in an email subject they depend on the encoder.
Compatibility
The list is Unicode’s own emoji test file, version 15.1 from 2023, and the names are the official CLDR short names rather than anything invented here. That matters for search: the name shown is the one your operating system also uses. Unicode data is used under the Unicode License.
1,432 of the file’s 3,782 sequences are in the picker. The ones left out are skin tone variants, which are the base plus a modifier and would make the grid six times longer; flags, of which there are 258 and which belong in their own list; gendered duplicates where a neutral form exists; and anything newer than Unicode 14, which still renders as a box on plenty of machines. The inspector works on all of them, and on anything else you paste.
Everything runs in your browser. Nothing is sent anywhere, and the grid is capped at 240 results at a time because laying out 1,432 buttons costs more than it is worth.
Grapheme counting uses Intl.Segmenter, which is in every current browser and falls back to counting
code points where it is missing, so the “characters as people see them” figure is the browser’s own
opinion rather than a guess.
Frequently asked questions
Why does my database reject emoji?
utf8 rather than utf8mb4. Converting the table and the connection
charset fixes it; on WordPress, DB_CHARSET in wp-config.php should be utf8mb4.How do I count characters correctly?
Array.from( text ).length gives code points, which is closer than .length and still not what a
person sees.