Duplicate Line Remover
Removes duplicate lines while keeping the order, and reports how many more duplicates trimming and case folding found than an exact comparison would.
3 [email protected]
2 [email protected]
1 [email protected]
----------------------------------------
Lines in 6
Blank lines 0
Unique entries 3
Duplicated entries 2
Lines removed 3
Lines out 3
Mode keep first
Compared trimmed, case ignored
Order as they first appeared
Trimming and case folding found 3 more duplicates than an exact
comparison would. That gap is the interesting number: it means the list
has entries that differ only in whitespace or capitalisation, which is
what a spreadsheet paste and a hand-typed list both produce.
Order is preserved: each entry stays where it first appeared, which is
usually why the list is in that order. Turn sorting on when you want to
scan for near-duplicates by eye.
The top line is not always the one you want to keep. "Keep last" exists
for the case where later entries are corrections of earlier ones, which
is what an append-only log or a form export looks like. Here that
changes 2 entries.
For a list of email addresses, remember that case matters in the local
part by the letter of the specification and no real provider enforces
it, while the domain is definitively case insensitive. Folding case is
the practical choice and it is not strictly correct.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Two decisions matter when removing duplicate lines, and both are usually made for you by whatever tool you reach for.
The first is order. Keeping the first occurrence preserves the order somebody put the list
in, which is generally the reason the list exists. Sorting makes duplicates adjacent and
throws that order away. sort -u does the second, and people reach for it expecting the
first.
The second is what counts as the same. “Apple” and “apple ” are different strings and usually the same entry. Trimming and case folding are options here rather than assumptions, and the report says how many extra duplicates each one found. That gap is the interesting number: it tells you whether the list is repetitive or dirty.
How to use
- Paste the list, one entry per line.
- Choose what to keep: the first occurrence, the last, only the duplicates, or only the entries that appear once.
- Read the summary before you use the output.
Example
Five email addresses, two of them duplicates only if you trim and fold case:
3 [email protected]
2 [email protected]
1 [email protected]
Lines in 6
Unique entries 3
Duplicated entries 2
Lines removed 3
Lines out 3
Mode keep first
Compared trimmed, case ignored
Order as they first appeared
An exact comparison would have found six unique entries and removed nothing. The three
removed here are [email protected] with a trailing space, [email protected] in a
different case, and [email protected] likewise, which is exactly what a spreadsheet paste
and a hand-typed list produce between them.
Pitfalls
sort -u loses your order. If the list is a priority order, a chronological log or a
column pasted from a spreadsheet, sorting destroys the thing that made it useful. Keep
sorting for when you want to scan by eye.
“Keep last” is not the same as “keep first”. It exists for lists where later entries correct earlier ones: an append-only log, a form export, a changelog. The entry stays in its original position and its text comes from the last occurrence, so the order survives and the value is updated.
Case folding email addresses is practical and not strictly correct. By the letter of the
specification the local part is case sensitive, so [email protected] and [email protected]
may be two mailboxes. No real provider treats them differently, and the domain is
definitively case insensitive. Fold case and know that you have made a judgement.
Trailing whitespace is the commonest hidden duplicate. It is invisible in every editor and it is what a paste from a table produces. Trimming is on by default for that reason.
“Only duplicates” is for finding, not removing. It shows the entries that appear more than once, which is what you want when the duplicates are the problem to investigate rather than the thing to delete.
Blank lines count as an entry unless you drop them. A list with three blank lines produces one blank entry with a count of three, which is correct and rarely what you meant.
This is not fuzzy matching. “Jon Smith” and “John Smith” are different entries, as are
example.com and www.example.com. Near-duplicate detection needs a similarity measure and
a threshold you would have to argue about, and guessing silently would be worse than not
guessing.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored, which matters for a list of customer emails.
The comparison is exact string equality on the key, after the trimming and case folding you
choose. Case folding uses JavaScript’s own toLowerCase, which is Unicode aware, so
accented characters fold correctly and the Turkish dotless i behaves as Unicode says rather
than as Turkish locale rules would.
Sorting uses a locale-aware comparison with numeric ordering on, so item2 comes before
item10 rather than after it, which is almost always what you wanted from a list with
numbers in it.
The counts column is tab separated, so the output pastes into a spreadsheet as two columns without any further work.
Input is capped at 100,000 characters. The work is linear; the limit is the browser holding the input and the output at the same time.
Frequently asked questions
How do I remove duplicates in Excel instead?
Can I keep the order but sort within duplicates?
Why is my count higher than the number of lines I can see?
What about duplicates that differ by a trailing slash or www?
Is there a size limit?
sort -u or awk '!seen[$0]++' at a terminal is the right tool,
and the second one preserves order.