Text Reverser

Reverses characters, words or lines by grapheme cluster, and shows what the split-reverse-join one-liner does to an emoji and a combining accent.

Enable JavaScript to customise; default output below.

Reverse

Characters reverses within each word as well. Words keeps each word intact and reverses their order.

Live preview reversed.txt
anañam 👨‍👩‍👧 éfac

Characters a reader sees  13
Unicode code points       17
UTF-16 code units         20

Three ways to reverse it
  by grapheme cluster     anañam 👨‍👩‍👧 éfac
  by code point           anañam 👧‍👩‍👨 éfac
  by UTF-16 code unit     anañam ��‍��‍�� éfac
  unpaired surrogates     6, drawn above as the replacement character

They disagree
  code units              split something into halves that are not characters
  and those halves        cannot be stored: JSON, MySQL and most APIs reject a lone surrogate
  code points             moved a combining mark or broke a joined sequence

Palindrome                no

`text.split( "" ).reverse().join( "" )` is the answer in every tutorial
and it is wrong. It splits on UTF-16 code units, so an emoji outside the
basic plane comes apart into its two surrogate halves, and the halves
are not characters: they arrive as replacement squares.

Splitting by code point, with the spread operator or `Array.from`, fixes
the emoji and not the accents. A combining acute is its own code point,
so reversing moves it onto whichever letter now precedes it, and "café"
written that way comes back with the accent in the wrong place.

What a reader calls a character is a grapheme cluster: a base character
plus its combining marks, a flag made of two regional indicators, a
family emoji joined with zero-width joiners. `Intl.Segmenter` with
grapheme granularity is the only method here that counts them, and it is
what the reversal above uses.

This text is 13 characters to a reader, 17 code points and 20 UTF-16
units. Where those three differ, so does every length check, every
truncation and every database column limit in the system.

A right-to-left script reverses correctly as clusters and still displays
confusingly, because the bidirectional algorithm lays out the result
according to its own rules. Reversal is a poor tool for Arabic or Hebrew
text and the answer will look wrong even when it is not.

Reversing is not encryption. It is not obfuscation either: a reversed
string is recognisable at a glance and reversible by anyone. If the
point is to hide something, this is the wrong tool and so is base64.

Output is valid and updates as you type.

text.split( '' ).reverse().join( '' ) is the answer in every tutorial about reversing a string, and it is wrong for any text that is not plain ASCII.

It splits on UTF-16 code units. An emoji outside the basic plane is two of those, so it comes apart into its surrogate halves, and the halves are not characters: they arrive as replacement squares. Splitting by code point instead fixes the emoji and not the accents, because a combining acute is its own code point and reversing moves it onto whichever letter now comes before it.

What a reader calls a character is a grapheme cluster. That is what this reverses.

How to use

  1. Paste the text.
  2. Pick characters, words or lines.
  3. For characters, read the three results: the right one and the two broken ones.

Example

café 👨‍👩‍👧 mañana, reversed:

anañam 👨‍👩‍👧 éfac

Characters a reader sees  13
Unicode code points       17
UTF-16 code units         20

Three ways to reverse it
  by grapheme cluster     anañam 👨‍👩‍👧 éfac
  by code point           anañam 👧‍👩‍👨 éfac
  by UTF-16 code unit     anañam ��‍��‍�� éfac
  unpaired surrogates     6, drawn above as the replacement character

They disagree
  code units              split something into halves that are not characters
  and those halves        cannot be stored: JSON, MySQL and most APIs reject a lone surrogate
  code points             moved a combining mark or broke a joined sequence

Palindrome                no

Thirteen characters, seventeen code points, twenty code units. Where those three numbers differ, so does every length check, truncation and column limit in the system.

Pitfalls

The one-liner breaks emoji. Any emoji above U+FFFF is a surrogate pair, and reversing code units swaps the two halves of it. The result is not a character and no font can draw it.

Splitting by code point is not enough. [ ...text ].reverse().join( '' ) keeps the emoji whole and still moves a combining mark to the wrong letter. It also reverses the members of a joined emoji: the family above becomes a different family, which at least renders, which is worse because you do not notice.

A flag is two characters that must stay in order. Regional indicator pairs make flags, so reversing the pair gives a different country or nothing at all. Grapheme clusters keep them together.

Right-to-left text reverses correctly and still looks wrong. Arabic and Hebrew are laid out by the bidirectional algorithm after the reversal, according to its own rules. Reversal is a poor tool for those scripts and the output will confuse a reader even when the operation was correct.

A lone surrogate cannot be stored. The halves the code-unit method leaves behind are not characters and are not valid Unicode, so JSON.stringify escapes them, PHP’s json_decode refuses them outright, and MySQL’s utf8mb4 will not accept them. A reversal that appears to work in the browser can therefore fail when the result is saved, which is a confusing bug to find. They are shown above as the replacement character, and counted.

Reversing is not encryption or obfuscation. A reversed string is recognisable at a glance and reversible by anyone. If the point is to hide something, this is the wrong tool, and so is base64.

Palindrome checking has the same problem. Comparing a string to its code-unit reverse fails on any text with an emoji or a combining accent in it. The check here folds case, drops punctuation and compares grapheme clusters.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored.

Grapheme clusters come from Intl.Segmenter, which is in every current browser: Chrome and Edge since 2020, Firefox since 2022, Safari 14.1 in 2021. Where it is missing there is a fallback that keeps a base character with its combining marks and keeps zero-width-joiner sequences together, which covers accents, flags and the joined emoji people actually paste, and is less complete than the real segmentation algorithm.

The test suite checks the three methods against a joined family emoji, a flag, and café written with a combining acute rather than a precomposed é, which is the case that catches code-point reversal.

In PHP, strrev() operates on bytes and will corrupt any multibyte character. There is no mb_strrev(); the usual workaround is preg_split( '//u', $text, -1, PREG_SPLIT_NO_EMPTY ), which splits by code point and therefore has the accent problem described above.

Frequently asked questions

Why does my reversed emoji show as two squares?
The code that reversed it split a surrogate pair. Each half is an unpaired surrogate, which is not a character, so the font draws the replacement glyph for both.
Is there a one-liner that is actually correct?
[ ...new Intl.Segmenter( undefined, { granularity: 'grapheme' } ).segment( s ) ].map( p => p.segment ).reverse().join( '' ). It is not short, and that is the honest length of the correct answer.
What is a grapheme cluster?
The unit a reader thinks of as one character: a letter plus its accents, a flag made of two regional indicators, a family emoji joined with zero-width joiners, an emoji plus a skin-tone modifier. Unicode defines the rules for finding the boundaries, and Intl.Segmenter implements them.
Does reversing words reverse the letters too?
No. The words mode keeps each word as it is and reverses their order, preserving the spacing between them. The characters mode reverses everything, which also reverses each word internally.
Can I use this to make upside-down text?
No, that is a different thing: it substitutes characters that look like inverted letters. The Unicode font generator on this site does that, and it comes with its own warning about screen readers.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.