Random Text Generator
Random test strings that actually break things: accents, emoji, combining marks, invisible characters, right-to-left runs and the empty cases.
generated in your browser
Strings 12
Categories accents and other scripts, emoji and joined sequences, quotes and apostrophes, strings that look like markup, long and unbroken, the empty cases
What each category is for
accents and other scripts one character can be several bytes, so a byte-length limit truncates mid-character
emoji and joined sequences these are two to eleven UTF-16 units, so anything counting characters by index breaks them
quotes and apostrophes the most common escaping failure there is, and half of them are not the ASCII ones
strings that look like markup these prove whether output is escaped; they are inert text on their own
long and unbroken no spaces means nowhere to wrap, which is how a layout overflows
the empty cases the ones a required-field check and a trim both have opinions about
Lorem ipsum is the wrong fixture for testing a field. It is plain ASCII
with tidy word lengths, so it passes every check and proves nothing. The
bugs are in the text it does not contain.
A byte limit is not a character limit. `varchar(20)` in a UTF-8 column
holds twenty bytes, so a name with accents fits fewer than twenty
characters, and truncating at a byte boundary can split a character in
half and produce invalid text.
The same visible string can be two different byte sequences. `café`
written with a precomposed é and with a combining acute look identical
and compare unequal, which is why search and deduplication need Unicode
normalisation rather than a plain comparison.
Anything that looks like markup or a query here is inert text. The test
is whether the thing receiving it escapes on output and parameterises
its queries; if it does, these strings are just strings. If they do
something, that is the finding.
Invisible characters survive copying. A zero-width space pasted into a
field makes two identical-looking values unequal, and a non-breaking
space is not caught by a trim that only looks for ASCII whitespace.
Mixed direction text reorders on display. A right-to-left run next to a
number or a Latin word is laid out by the bidirectional algorithm, so a
label and its value can appear to swap places even though the data is
correct.
The empty cases are worth more than the exotic ones. An empty string, a
single space, an ideographic space, the text "null", and the number zero
are the inputs that reveal what a required-field check and a falsy test
actually do.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Lorem ipsum is the wrong text to test a field with. It is plain ASCII, tidily worded, with no accents, no emoji, no combining marks, no invisible characters and no quotes. It passes everything, which is exactly the problem: the bugs are in the text it does not contain.
The name with an accent that breaks a byte-length limit. The emoji that arrives as two replacement squares. The apostrophe in O’Brien that ends a SQL string. The Arabic that reorders a table cell. The zero-width space that makes two identical-looking values compare unequal. The empty string, which half of every validation layer has a different opinion about.
So this generates those instead.
How to use
- Tick the categories you want to test against.
- Pick how many strings.
- Paste them into the field, one at a time, and watch what breaks.
Example
Twelve strings from the default categories:
1 👩💻 profession
2 Łukasz Dąbrowski
3 {{template}}
4 3️⃣ keycap
5 Håkon Wium Lie
6 🇧🇩 flag
7 Ñoño Güell
8 (1 invisible character)
contains U+3000 ideographic space
9 null
10 back`tick
11 Donaudampfschifffahrtselektrizitaetenhauptbetriebswerkbauunterbeamtengesellschaft
12 https://example.com/a/very/long/path/that/will/not/wrap/anywhere/at/all/index.html
The invisible ones are described rather than printed blank, because a row showing nothing is not a fixture you can read.
Pitfalls
A byte limit is not a character limit. varchar(20) in a UTF-8 column holds twenty bytes, so
Đặng Thị Hương needs more of them than it has characters. Truncating at a byte boundary can split a
character in half and produce text that is not valid UTF-8 at all.
The same visible string can be two byte sequences. café with a precomposed é and café with a
combining acute look identical and compare unequal. Search, deduplication and uniqueness constraints
all need Unicode normalisation, usually NFC, rather than a plain comparison.
The strings that look like code are inert. <script>ignored()</script> and 1' OR '1'='1 are text.
They test whether the receiving system escapes on output and parameterises its queries. If nothing
happens, that is the pass; if something happens, that is the finding. Only use them against systems you
are responsible for.
Emoji are not one character to a computer. A family emoji is eleven UTF-16 units joined with
zero-width joiners, so any code counting or slicing by index will break it. "👨👩👧".length is 8.
Invisible characters survive a copy and paste. A non-breaking space is not caught by a trim that looks for ASCII whitespace, a soft hyphen disappears in some fonts and not others, and a zero-width space makes two visually identical usernames different accounts.
Mixed-direction text reorders on display. A right-to-left run next to a number or a Latin word is laid out by the bidirectional algorithm, so a label and its value can appear to swap places even though the stored data is correct. That is a rendering fact, not a data bug, and it still confuses users.
The empty cases are worth more than the exotic ones. An empty string, one space, the word “null”, and the number zero are the inputs that reveal what a required-field check and a falsy test actually do, and they are the ones a manual tester skips.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored.
The samples are fixed lists chosen because each one has broken real software, and the selection from
them uses crypto.getRandomValues so a second run gives a different set. The test suite asserts that
each category really contains what it claims: that the accented samples take more UTF-8 bytes than they
have characters, that the emoji take more UTF-16 units, that the combining pair differs before
normalisation and matches after it, and that the long samples contain no space to wrap at.
Invisible characters are named by code point in the output, and a whitespace-only value is described rather than printed, so what you are looking at is unambiguous.
At build time there is no randomness, so the page ships with a placeholder and fills in when it loads.
Frequently asked questions
What should I test a text field with?
Is it safe to paste the SQL-looking strings?
What is Unicode normalisation?
String.prototype.normalize( 'NFC' )
in JavaScript, Normalizer::normalize in PHP. Do it on input if you care whether two names are the
same name.Why does my form strip the emoji?
utf8 rather than utf8mb4. The older one holds three bytes a
character and an emoji needs four, so it is truncated or rejected.