Text Reverser
Reverses characters, words or lines by grapheme cluster, and shows what the split-reverse-join one-liner does to an emoji and a combining accent.
anañam 👨👩👧 éfac
Characters a reader sees 13
Unicode code points 17
UTF-16 code units 20
Three ways to reverse it
by grapheme cluster anañam 👨👩👧 éfac
by code point anañam 👧👩👨 éfac
by UTF-16 code unit anañam ������ éfac
unpaired surrogates 6, drawn above as the replacement character
They disagree
code units split something into halves that are not characters
and those halves cannot be stored: JSON, MySQL and most APIs reject a lone surrogate
code points moved a combining mark or broke a joined sequence
Palindrome no
`text.split( "" ).reverse().join( "" )` is the answer in every tutorial
and it is wrong. It splits on UTF-16 code units, so an emoji outside the
basic plane comes apart into its two surrogate halves, and the halves
are not characters: they arrive as replacement squares.
Splitting by code point, with the spread operator or `Array.from`, fixes
the emoji and not the accents. A combining acute is its own code point,
so reversing moves it onto whichever letter now precedes it, and "café"
written that way comes back with the accent in the wrong place.
What a reader calls a character is a grapheme cluster: a base character
plus its combining marks, a flag made of two regional indicators, a
family emoji joined with zero-width joiners. `Intl.Segmenter` with
grapheme granularity is the only method here that counts them, and it is
what the reversal above uses.
This text is 13 characters to a reader, 17 code points and 20 UTF-16
units. Where those three differ, so does every length check, every
truncation and every database column limit in the system.
A right-to-left script reverses correctly as clusters and still displays
confusingly, because the bidirectional algorithm lays out the result
according to its own rules. Reversal is a poor tool for Arabic or Hebrew
text and the answer will look wrong even when it is not.
Reversing is not encryption. It is not obfuscation either: a reversed
string is recognisable at a glance and reversible by anyone. If the
point is to hide something, this is the wrong tool and so is base64.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
text.split( '' ).reverse().join( '' ) is the answer in every tutorial about reversing a string,
and it is wrong for any text that is not plain ASCII.
It splits on UTF-16 code units. An emoji outside the basic plane is two of those, so it comes apart into its surrogate halves, and the halves are not characters: they arrive as replacement squares. Splitting by code point instead fixes the emoji and not the accents, because a combining acute is its own code point and reversing moves it onto whichever letter now comes before it.
What a reader calls a character is a grapheme cluster. That is what this reverses.
How to use
- Paste the text.
- Pick characters, words or lines.
- For characters, read the three results: the right one and the two broken ones.
Example
café 👨👩👧 mañana, reversed:
anañam 👨👩👧 éfac
Characters a reader sees 13
Unicode code points 17
UTF-16 code units 20
Three ways to reverse it
by grapheme cluster anañam 👨👩👧 éfac
by code point anañam 👧👩👨 éfac
by UTF-16 code unit anañam ������ éfac
unpaired surrogates 6, drawn above as the replacement character
They disagree
code units split something into halves that are not characters
and those halves cannot be stored: JSON, MySQL and most APIs reject a lone surrogate
code points moved a combining mark or broke a joined sequence
Palindrome no
Thirteen characters, seventeen code points, twenty code units. Where those three numbers differ, so does every length check, truncation and column limit in the system.
Pitfalls
The one-liner breaks emoji. Any emoji above U+FFFF is a surrogate pair, and reversing code units swaps the two halves of it. The result is not a character and no font can draw it.
Splitting by code point is not enough. [ ...text ].reverse().join( '' ) keeps the emoji whole
and still moves a combining mark to the wrong letter. It also reverses the members of a joined
emoji: the family above becomes a different family, which at least renders, which is worse because
you do not notice.
A flag is two characters that must stay in order. Regional indicator pairs make flags, so reversing the pair gives a different country or nothing at all. Grapheme clusters keep them together.
Right-to-left text reverses correctly and still looks wrong. Arabic and Hebrew are laid out by the bidirectional algorithm after the reversal, according to its own rules. Reversal is a poor tool for those scripts and the output will confuse a reader even when the operation was correct.
A lone surrogate cannot be stored. The halves the code-unit method leaves behind are not
characters and are not valid Unicode, so JSON.stringify escapes them, PHP’s json_decode refuses
them outright, and MySQL’s utf8mb4 will not accept them. A reversal that appears to work in the
browser can therefore fail when the result is saved, which is a confusing bug to find. They are shown
above as the replacement character, and counted.
Reversing is not encryption or obfuscation. A reversed string is recognisable at a glance and reversible by anyone. If the point is to hide something, this is the wrong tool, and so is base64.
Palindrome checking has the same problem. Comparing a string to its code-unit reverse fails on any text with an emoji or a combining accent in it. The check here folds case, drops punctuation and compares grapheme clusters.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored.
Grapheme clusters come from Intl.Segmenter, which is in every current browser: Chrome and Edge
since 2020, Firefox since 2022, Safari 14.1 in 2021. Where it is missing there is a fallback that
keeps a base character with its combining marks and keeps zero-width-joiner sequences together,
which covers accents, flags and the joined emoji people actually paste, and is less complete than
the real segmentation algorithm.
The test suite checks the three methods against a joined family emoji, a flag, and café written
with a combining acute rather than a precomposed é, which is the case that catches code-point
reversal.
In PHP, strrev() operates on bytes and will corrupt any multibyte character. There is no
mb_strrev(); the usual workaround is preg_split( '//u', $text, -1, PREG_SPLIT_NO_EMPTY ), which
splits by code point and therefore has the accent problem described above.
Frequently asked questions
Why does my reversed emoji show as two squares?
Is there a one-liner that is actually correct?
[ ...new Intl.Segmenter( undefined, { granularity: 'grapheme' } ).segment( s ) ].map( p => p.segment ).reverse().join( '' ).
It is not short, and that is the honest length of the correct answer.What is a grapheme cluster?
Intl.Segmenter implements
them.