Text Cleaner

Finds the characters you cannot see, counts them before removing any: zero-width spaces, non-breaking spaces, soft hyphens and BOMs.

Enable JavaScript to customise; default output below.

Paste what came out of Word, a PDF or a chat window. The default text has real invisible characters in it, so you can see the report work.

Live preview clean.txt
Helloworld and a softhyphen
Smart “quotes” and an em—dash…
Line with trailing spaces

Too many blank lines

----------------------------------------
Characters in                       113
Characters out                      106
Removed                             7

Found
  zero-width space                  1 — breaks a search for the word it sits inside
  soft hyphen                       1 — splits a word for a search engine and not for a reader
  non-breaking space                1 — breaks selectors, code and CSV parsing
  smart quotes, dashes or ellipses  4 — correct in prose, wrong in code and in a CSV
  lines with trailing whitespace    1 — invisible, and noise in every diff

The counts are taken before anything is changed, because knowing what
was in there is the useful part. A zero-width space or a non-breaking
space in a product code, a class name or a CSV column is invisible in
every editor and breaks the thing that reads it.

Smart quotes were left alone, which is right for prose. Turn the option
on when the text is going into code, a CSV file or anywhere a parser
will read it.

Line endings are normalised to \n whatever you started with, because a
file with both CRLF and LF in it is a file that shows a change on every
line of the next diff.

What this cannot fix: a homoglyph. A Cyrillic "а" looks exactly like a
Latin "a" and is a different character, so it survives cleaning and
still breaks a search. If a string looks right and behaves wrong, that
is the thing to suspect next.

Output is valid and updates as you type.

The characters that cause trouble are the ones you cannot see.

A zero-width space inside a product code breaks the search that would have found it. A non-breaking space in a CSS class name breaks the selector. A soft hyphen copied out of a PDF splits a word for a search engine and not for a reader. A byte order mark at the top of a file shows up as  the first time something reads it as the wrong encoding.

So this counts them first, before it changes anything, and names what each one breaks. The count is the useful part: knowing there were four non-breaking spaces in a column heading explains a bug that otherwise looks impossible.

How to use

  1. Paste the text. The default sample has real invisible characters in it, so you can see the report work.
  2. Read what was found.
  3. Turn off anything you want to keep. Smart quotes are off by default because they are correct in prose.

Example

Characters in                       118
Characters out                      109
Removed                             9

Found
  zero-width space                  1 — breaks a search for the word it sits inside
  soft hyphen                       1 — splits a word for a search engine and not for a reader
  non-breaking space                1 — breaks selectors, code and CSV parsing
  smart quotes, dashes or ellipses  4 — correct in prose, wrong in code and in a CSV
  lines with trailing whitespace    2 — invisible, and noise in every diff

Nine characters removed out of 118, and every one of them was invisible in the editor the text came from.

Pitfalls

Smart quotes are not dirt. Curly quotes, a real em dash and a proper ellipsis are correct typography, and straightening them in body copy makes the writing worse. The reason to straighten is destination: code, a CSV file, a URL, a password field, a JSON string. That is why the option is off by default.

A non-breaking space is the most expensive invisible character. Word inserts them, and they arrive in class names, in slugs, in spreadsheet cells and in code samples. Everything that looks for a space fails to find one.

A soft hyphen is legitimate in the wrong place. It marks where a word may break, which is a real typographic tool and a disaster in a database value. Copied from a justified PDF, it lands in the middle of words and survives every visual check.

Zero-width joiners are sometimes deliberate. They build emoji sequences and they carry meaning in Persian and Hindi text. Removing them from an emoji breaks it into its components. If the text is not English prose, look at the report before you clean.

Homoglyphs survive. A Cyrillic “а” looks exactly like a Latin “a” and is a different character. Nothing here catches it, because there is no way to tell an intentional Cyrillic word from an accidental one. If a string looks right and behaves wrong after cleaning, that is the thing to suspect next.

Trailing whitespace is noise in every diff. It is invisible and it changes the line, so a file with it shows edits on lines nobody touched. Removing it is free and it is on by default.

Collapsing spaces leaves indentation alone. A run at the start of a line is usually deliberate, in code and in poetry. A run in the middle of a sentence almost never is.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored, which matters because the text you are cleaning is often something you cannot paste into a website you do not control.

The nine invisible characters are the ones that turn up in real pasted text: zero-width space, joiner and non-joiner, byte order mark, soft hyphen, the left-to-right and right-to-left marks, the word joiner and the Mongolian vowel separator. The seven odd spaces are non-breaking, figure, thin, narrow no-break, ideographic, and the line and paragraph separators, which become newlines rather than spaces because that is what they mean.

Line endings are normalised to \n whatever you started with. A file with both CRLF and LF in it shows a change on every line of the next diff, which is the reason to do this even when nothing else needs doing.

Counts are taken from the original text, so the report describes what you pasted rather than what came out. That ordering is deliberate: a report on the cleaned text would always say zero.

Frequently asked questions

How did a zero-width space get into my text?
Chat clients, some CMS editors, PDF copy and paste, and anything that has been through a line-breaking algorithm. It is also occasionally added on purpose, to defeat a naive spam filter or to break up a URL.
Why is my CSV failing to parse?
Check for non-breaking spaces and for smart quotes. A quoted field delimited with curly quotes is not a quoted field to a parser, and a non-breaking space is not whitespace.
Will this fix encoding problems?
No. é where you expected é is a file read as the wrong encoding, and the fix is at the read rather than in the text. What this does catch is the byte order mark that often comes with it.
Should I clean before or after translating?
After. Translation tools sometimes introduce their own directional marks, which you would otherwise be removing and reintroducing.
Can I see which character is where?
Not in this tool. For that, an editor with “show invisibles” turned on, or a hex view, will point at the exact byte. The counts here tell you what to look for.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.