Random Text Generator

Random test strings that actually break things: accents, emoji, combining marks, invisible characters, right-to-left runs and the empty cases.

Live output

Enable JavaScript to customise; default output below.

Live preview test-strings.txt
generated in your browser

Strings                          12
Categories                       accents and other scripts, emoji and joined sequences, quotes and apostrophes, strings that look like markup, long and unbroken, the empty cases

What each category is for
  accents and other scripts      one character can be several bytes, so a byte-length limit truncates mid-character
  emoji and joined sequences     these are two to eleven UTF-16 units, so anything counting characters by index breaks them
  quotes and apostrophes         the most common escaping failure there is, and half of them are not the ASCII ones
  strings that look like markup  these prove whether output is escaped; they are inert text on their own
  long and unbroken              no spaces means nowhere to wrap, which is how a layout overflows
  the empty cases                the ones a required-field check and a trim both have opinions about

Lorem ipsum is the wrong fixture for testing a field. It is plain ASCII
with tidy word lengths, so it passes every check and proves nothing. The
bugs are in the text it does not contain.

A byte limit is not a character limit. `varchar(20)` in a UTF-8 column
holds twenty bytes, so a name with accents fits fewer than twenty
characters, and truncating at a byte boundary can split a character in
half and produce invalid text.

The same visible string can be two different byte sequences. `café`
written with a precomposed é and with a combining acute look identical
and compare unequal, which is why search and deduplication need Unicode
normalisation rather than a plain comparison.

Anything that looks like markup or a query here is inert text. The test
is whether the thing receiving it escapes on output and parameterises
its queries; if it does, these strings are just strings. If they do
something, that is the finding.

Invisible characters survive copying. A zero-width space pasted into a
field makes two identical-looking values unequal, and a non-breaking
space is not caught by a trim that only looks for ASCII whitespace.

Mixed direction text reorders on display. A right-to-left run next to a
number or a Latin word is laid out by the bidirectional algorithm, so a
label and its value can appear to swap places even though the data is
correct.

The empty cases are worth more than the exotic ones. An empty string, a
single space, an ideographic space, the text "null", and the number zero
are the inputs that reveal what a required-field check and a falsy test
actually do.

Output is valid and updates as you type.

Lorem ipsum is the wrong text to test a field with. It is plain ASCII, tidily worded, with no accents, no emoji, no combining marks, no invisible characters and no quotes. It passes everything, which is exactly the problem: the bugs are in the text it does not contain.

The name with an accent that breaks a byte-length limit. The emoji that arrives as two replacement squares. The apostrophe in O’Brien that ends a SQL string. The Arabic that reorders a table cell. The zero-width space that makes two identical-looking values compare unequal. The empty string, which half of every validation layer has a different opinion about.

So this generates those instead.

How to use

  1. Tick the categories you want to test against.
  2. Pick how many strings.
  3. Paste them into the field, one at a time, and watch what breaks.

Example

Twelve strings from the default categories:

1   👩‍💻 profession
2   Łukasz Dąbrowski
3   {{template}}
4   3️⃣ keycap
5   Håkon Wium Lie
6   🇧🇩 flag
7   Ñoño Güell
8   (1 invisible character)
      contains  U+3000 ideographic space
9   null
10  back`tick
11  Donaudampfschifffahrtselektrizitaetenhauptbetriebswerkbauunterbeamtengesellschaft
12  https://example.com/a/very/long/path/that/will/not/wrap/anywhere/at/all/index.html

The invisible ones are described rather than printed blank, because a row showing nothing is not a fixture you can read.

Pitfalls

A byte limit is not a character limit. varchar(20) in a UTF-8 column holds twenty bytes, so Đặng Thị Hương needs more of them than it has characters. Truncating at a byte boundary can split a character in half and produce text that is not valid UTF-8 at all.

The same visible string can be two byte sequences. café with a precomposed é and café with a combining acute look identical and compare unequal. Search, deduplication and uniqueness constraints all need Unicode normalisation, usually NFC, rather than a plain comparison.

The strings that look like code are inert. <script>ignored()</script> and 1' OR '1'='1 are text. They test whether the receiving system escapes on output and parameterises its queries. If nothing happens, that is the pass; if something happens, that is the finding. Only use them against systems you are responsible for.

Emoji are not one character to a computer. A family emoji is eleven UTF-16 units joined with zero-width joiners, so any code counting or slicing by index will break it. "👨‍👩‍👧".length is 8.

Invisible characters survive a copy and paste. A non-breaking space is not caught by a trim that looks for ASCII whitespace, a soft hyphen disappears in some fonts and not others, and a zero-width space makes two visually identical usernames different accounts.

Mixed-direction text reorders on display. A right-to-left run next to a number or a Latin word is laid out by the bidirectional algorithm, so a label and its value can appear to swap places even though the stored data is correct. That is a rendering fact, not a data bug, and it still confuses users.

The empty cases are worth more than the exotic ones. An empty string, one space, the word “null”, and the number zero are the inputs that reveal what a required-field check and a falsy test actually do, and they are the ones a manual tester skips.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored.

The samples are fixed lists chosen because each one has broken real software, and the selection from them uses crypto.getRandomValues so a second run gives a different set. The test suite asserts that each category really contains what it claims: that the accented samples take more UTF-8 bytes than they have characters, that the emoji take more UTF-16 units, that the combining pair differs before normalisation and matches after it, and that the long samples contain no space to wrap at.

Invisible characters are named by code point in the output, and a whitespace-only value is described rather than printed, so what you are looking at is unambiguous.

At build time there is no randomness, so the page ships with a placeholder and fills in when it loads.

Frequently asked questions

What should I test a text field with?
One from each category here, plus the longest value the field claims to accept and one character more. The categories are ordered roughly by how often they find something.
Is it safe to paste the SQL-looking strings?
Into your own application, yes: they are text, and a parameterised query treats them as text. Do not use them against systems you are not responsible for, which is a question about permission rather than about the strings.
What is Unicode normalisation?
Turning equivalent sequences into one canonical form so they compare equal. String.prototype.normalize( 'NFC' ) in JavaScript, Normalizer::normalize in PHP. Do it on input if you care whether two names are the same name.
Why does my form strip the emoji?
Almost always a MySQL column using utf8 rather than utf8mb4. The older one holds three bytes a character and an emoji needs four, so it is truncated or rejected.
Do you have a name generator or fake addresses?
No. This is about the characters, not the content. For realistic fake data a dedicated library is the better tool, and it will generate mostly ASCII, which is why these categories are still worth a separate pass.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.