Duplicate Line Remover

Removes duplicate lines while keeping the order, and reports how many more duplicates trimming and case folding found than an exact comparison would.

Enable JavaScript to customise; default output below.

Keep

"Keep last" is for a list where later entries correct earlier ones. "Only duplicates" is for finding them rather than removing them.

Live preview lines.txt
3	[email protected]
2	[email protected]
1	[email protected]

----------------------------------------
Lines in            6
Blank lines         0
Unique entries      3
Duplicated entries  2
Lines removed       3
Lines out           3

Mode                keep first
Compared            trimmed, case ignored
Order               as they first appeared

Trimming and case folding found 3 more duplicates than an exact
comparison would. That gap is the interesting number: it means the list
has entries that differ only in whitespace or capitalisation, which is
what a spreadsheet paste and a hand-typed list both produce.

Order is preserved: each entry stays where it first appeared, which is
usually why the list is in that order. Turn sorting on when you want to
scan for near-duplicates by eye.

The top line is not always the one you want to keep. "Keep last" exists
for the case where later entries are corrections of earlier ones, which
is what an append-only log or a form export looks like. Here that
changes 2 entries.

For a list of email addresses, remember that case matters in the local
part by the letter of the specification and no real provider enforces
it, while the domain is definitively case insensitive. Folding case is
the practical choice and it is not strictly correct.

Output is valid and updates as you type.

Two decisions matter when removing duplicate lines, and both are usually made for you by whatever tool you reach for.

The first is order. Keeping the first occurrence preserves the order somebody put the list in, which is generally the reason the list exists. Sorting makes duplicates adjacent and throws that order away. sort -u does the second, and people reach for it expecting the first.

The second is what counts as the same. “Apple” and “apple ” are different strings and usually the same entry. Trimming and case folding are options here rather than assumptions, and the report says how many extra duplicates each one found. That gap is the interesting number: it tells you whether the list is repetitive or dirty.

How to use

  1. Paste the list, one entry per line.
  2. Choose what to keep: the first occurrence, the last, only the duplicates, or only the entries that appear once.
  3. Read the summary before you use the output.

Example

Five email addresses, two of them duplicates only if you trim and fold case:

3	[email protected]
2	[email protected]
1	[email protected]

Lines in            6
Unique entries      3
Duplicated entries  2
Lines removed       3
Lines out           3

Mode                keep first
Compared            trimmed, case ignored
Order               as they first appeared

An exact comparison would have found six unique entries and removed nothing. The three removed here are [email protected] with a trailing space, [email protected] in a different case, and [email protected] likewise, which is exactly what a spreadsheet paste and a hand-typed list produce between them.

Pitfalls

sort -u loses your order. If the list is a priority order, a chronological log or a column pasted from a spreadsheet, sorting destroys the thing that made it useful. Keep sorting for when you want to scan by eye.

“Keep last” is not the same as “keep first”. It exists for lists where later entries correct earlier ones: an append-only log, a form export, a changelog. The entry stays in its original position and its text comes from the last occurrence, so the order survives and the value is updated.

Case folding email addresses is practical and not strictly correct. By the letter of the specification the local part is case sensitive, so [email protected] and [email protected] may be two mailboxes. No real provider treats them differently, and the domain is definitively case insensitive. Fold case and know that you have made a judgement.

Trailing whitespace is the commonest hidden duplicate. It is invisible in every editor and it is what a paste from a table produces. Trimming is on by default for that reason.

“Only duplicates” is for finding, not removing. It shows the entries that appear more than once, which is what you want when the duplicates are the problem to investigate rather than the thing to delete.

Blank lines count as an entry unless you drop them. A list with three blank lines produces one blank entry with a count of three, which is correct and rarely what you meant.

This is not fuzzy matching. “Jon Smith” and “John Smith” are different entries, as are example.com and www.example.com. Near-duplicate detection needs a similarity measure and a threshold you would have to argue about, and guessing silently would be worse than not guessing.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored, which matters for a list of customer emails.

The comparison is exact string equality on the key, after the trimming and case folding you choose. Case folding uses JavaScript’s own toLowerCase, which is Unicode aware, so accented characters fold correctly and the Turkish dotless i behaves as Unicode says rather than as Turkish locale rules would.

Sorting uses a locale-aware comparison with numeric ordering on, so item2 comes before item10 rather than after it, which is almost always what you wanted from a list with numbers in it.

The counts column is tab separated, so the output pastes into a spreadsheet as two columns without any further work.

Input is capped at 100,000 characters. The work is linear; the limit is the browser holding the input and the output at the same time.

Frequently asked questions

How do I remove duplicates in Excel instead?
Data, then Remove Duplicates. It keeps the first occurrence and it does not trim or fold case, which is why a spreadsheet often reports fewer duplicates than this does.
Can I keep the order but sort within duplicates?
That is what “keep first” does: the entry appears where it first appeared. If you want grouped duplicates as well, run it with “only duplicates” to get that list separately.
Why is my count higher than the number of lines I can see?
Blank lines, or trailing whitespace making two lines that look identical into one entry with a count of two. The summary separates blanks out for exactly this reason.
What about duplicates that differ by a trailing slash or www?
Those are different strings and this will keep both. Normalising URLs before comparing is a separate job, and doing it invisibly here would be the wrong kind of helpful.
Is there a size limit?
100,000 characters, which is roughly two thousand email addresses or five thousand short lines. For a bigger file, sort -u or awk '!seen[$0]++' at a terminal is the right tool, and the second one preserves order.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.