Text Merger
Merges two lists by pairing, alternating or appending, normalising the line endings first and saying what happened to the rows with no partner.
apple: red
banana: green
cherry: dark red
damson: (no colour)
Left list 4 rows
Right list 3 rows
Out 4 rows
Unpaired rows 1, all on the left
what happened to them paired with "(no colour)"
Line endings
left LF
right LF
The carriage return is the bug you cannot see. A list pasted from
Windows has a `\r` on the end of every line, so joining it to a Unix
list puts that character in the middle of the row. `"apple\r"` and
`"apple"` look identical and compare unequal, which then breaks a
lookup, a sort or a deduplication somewhere downstream. Both lists are
normalised here before anything else happens.
There is 1 unpaired row on the left, which is the thing to decide about
deliberately. Stopping at the shorter list silently loses data, so it is
worth knowing that is what happened.
A trailing newline leaves an empty last element that is not a row.
Pairing it with a real value from the other list shifts every row below
it by one, and the result looks plausible, which is the dangerous part.
Blank rows are dropped here before pairing for that reason.
A merge is not a join. Pairing by position assumes the two lists are
already in the same order, and nothing here checks that. If the rows
have a key, join on the key in a spreadsheet or a database instead; if
they do not, verify the order before trusting the output.
Deduplicating after the merge removes identical joined rows, not
duplicate values in one column. Two rows that differ only by trailing
whitespace are different rows, which is another reason the normalisation
above matters.
For anything with commas in the values, a comma separator produces a
file that is not valid CSV. Either quote the values or use a tab, and a
tab-separated file is the safer choice for something going into a
spreadsheet.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Two lists, one pasted from a Windows machine and one from a Mac, merged into pairs. The output looks right. Then a lookup against those rows fails, a sort puts them in a strange order, and a deduplication finds nothing.
The reason is a carriage return. A Windows list carries a \r on the end of every line, so joining it
to another list puts that character in the middle of the row: "apple\r" and "apple" look identical
and are not equal. Both lists are normalised here before anything else happens, and the mismatch is
reported rather than silently absorbed.
The other decision worth making on purpose is what happens to the rows with no partner.
How to use
- Paste the two lists, one row a line.
- Choose how to merge: pair them, alternate them, append one to the other, or wrap every row of the first in the first two lines of the second.
- Decide what to do with the rows that have no partner.
Example
Four fruits against three colours, padding the leftover:
apple: red
banana: green
cherry: dark red
damson: (no colour)
Left list 4 rows
Right list 3 rows
Out 4 rows
Unpaired rows 1, all on the left
what happened to them paired with "(no colour)"
Line endings
left CRLF
right LF
note the two lists disagree, and both were normalised before merging
Stopping at the shorter list would have produced three tidy rows and lost the damson without saying so.
Pitfalls
The carriage return is invisible and it breaks equality. It survives a copy from a Windows editor, a
CSV exported by Excel, and anything that came through a Windows-hosted FTP. It is normalised here; in
your own code, split on /\r\n|\r|\n/ rather than on '\n'.
Stopping at the shorter list loses data silently. It is often the right choice and it should be a choice. The count of unpaired rows and which side they are on is printed for that reason.
A trailing newline leaves an empty last element. Pairing it with a real value from the other list shifts every row below it by one, and the result is plausible, which is the dangerous part. Blank rows are dropped before pairing here.
A merge is not a join. Pairing by position assumes the two lists are already in the same order, and nothing checks that. If the rows have a key, join on the key in a spreadsheet or a database. If they do not, verify the order before trusting the output.
A comma separator does not make a CSV. Any value containing a comma breaks the file. Use a tab, or quote the values properly; a tab-separated file pastes into a spreadsheet more reliably than a hand-built CSV.
Deduplicating removes identical joined rows. It does not find duplicate values in one column, and two rows differing only in trailing whitespace are two different rows, which is another reason the normalisation matters.
Alternating and pairing are different operations. Alternating produces one row per input row from both lists, interleaved. Pairing produces one row per pair. Picking the wrong one is obvious in the output and easy to do when the lists are long.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored, which matters when the two lists are customer data.
All three line endings are recognised and normalised, CRLF is counted once rather than as a CR plus an LF, and the dominant style of each list is reported so a mismatch is visible. The test suite asserts that no output row contains a carriage return when one list arrives with them.
The prefix-suffix mode takes the first two lines of the second list as the wrapper, which is the quickest
way to turn a list into markup: <li> and </li> produce a list, " and ", produce a JSON-ish array
that still needs escaping checked.
Blank rows are dropped before pairing, and the count of what went in and what came out is always printed so a row that disappeared is visible in the numbers.
Frequently asked questions
How do I combine two lists side by side?
What if the lists are in different orders?
Can I merge more than two lists?
How do I wrap every line in quotes?
" on the first line of the second box and ", on the second. Check the result
if any value contains a quote of its own, because that needs escaping and this does not do it for you.