HTML Merger
Merges several HTML documents into one, de-duplicating head elements and reporting every drop with the document it came from.
<!doctype html>
<html lang="en">
<head>
<title>Shop</title>
<meta charset="utf-8" />
<link rel="stylesheet" href="/app.css" />
<link rel="stylesheet" href="/main.css" />
</head>
<body>
<!-- header.html -->
<header>The header</header>
<!-- main.html -->
<main>The main content</main>
<!-- document 3 -->
<section>A partial with no head or body of its own</section>
</body>
</html>
Documents 3
header.html 181 characters
main.html 204 characters
document 3 60 characters
Fragments 1 with no head or body of their own
Dropped 3
"Shop, main" from main.html was dropped: a document has one title
a duplicate charset from main.html was dropped
a duplicate stylesheet or link from main.html was dropped
3 elements were removed as duplicates, and each one is listed above with
the document it came from. A merge that drops things silently is the
reason this list exists: the second copy of a stylesheet is usually
harmless and occasionally the one with the override in it.
The parsing is the same parser the rest of these tools use, so a
"</body>" inside a script string is not a closing tag and an unclosed
div in a partial does not swallow the next document. That is the
difference between a merge and a concatenation.
Head order is kept as found, first document first, which matters for
stylesheets: two rules of equal specificity are decided by which came
last. If the merged page looks wrong, that ordering is the first thing
to check.
Scripts are moved as they are. An inline script that expects to run
before an element exists will still run before it, and one that depended
on being the only script on the page is now not.
Inline styles from two documents can collide even when neither file
changes. The same class name in two partials becomes one cascade here,
which the browser resolves by order rather than by intent.
A charset declaration has to be in the first 1024 bytes of the document
to be honoured, which it is here because the head is written before
anything else. It is worth knowing if you reorder the result by hand.
This produces one document, not a build step. For anything that has to
happen repeatedly, the same merge belongs in a template or a bundler,
where the duplicate handling is a configuration rather than a report.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Merging HTML is two decisions: what goes in the head once, and what order the bodies go in.
The first is where a merge goes wrong. Two documents that each load the same stylesheet should load it once, two that each declare a charset should declare it once, and two that each have a title have to lose one. A tool that drops those silently is indistinguishable from a tool that drops the wrong one, so everything removed here is listed with the document it came from.
The parsing is the parser these tools already use rather than a second set of regular expressions. That
matters in practice: a </body> inside a script string is not a closing tag, and an unclosed <div> in a
pasted partial does not swallow the next document.
How to use
- Paste the documents, separated by a line of three or more dashes.
- Name each one with a comment like
<!-- file: header.html -->. The report refers to them by that name. - Read the dropped list before using the result. It is the part worth checking.
Example
Two pages and a partial:
<!doctype html>
<html lang="en">
<head>
<title>Shop</title>
<meta charset="utf-8" />
<link rel="stylesheet" href="/app.css" />
<link rel="stylesheet" href="/main.css" />
</head>
<body>
<!-- header.html -->
<header>The header</header>
<!-- main.html -->
<main>The main content</main>
<!-- document 3 -->
<section>A partial with no head or body of its own</section>
</body>
</html>
And the report underneath it:
Documents 3
header.html ...
main.html ...
document 3 ...
Fragments 1 with no head or body of their own
Dropped 3
"Shop, main" from main.html was dropped: a document has one title
a duplicate charset from main.html was dropped
a duplicate stylesheet or link from main.html was dropped
Pitfalls
Read the dropped list. The second copy of a stylesheet is usually harmless and occasionally the one with the override in it.
Head order decides the cascade. Elements are kept in the order found, first document first. Two rules of equal specificity are resolved by which came last, so if the merged page looks wrong, start there.
Scripts move as they are. An inline script that ran before an element existed still does, and one that assumed it was the only script on the page no longer is.
Class names collide. Two partials with the same class name become one cascade here. Neither file changed; the context did.
A separator is a line of its own. Three dashes inside a line of markup are not a separator. If the report says every document was a fragment, the separator did not match.
This is not a build step. For something that happens repeatedly, the same merge belongs in a template or a bundler, where duplicate handling is configuration rather than a report to read.
Charset has to be early. It is honoured only in the first 1024 bytes, which it is here because the head is written first. Worth remembering if you reorder the result by hand.
Compatibility
Runs in the browser: nothing is uploaded and nothing is stored, which matters for a page you have not published yet.
The merge uses the project’s own HTML parser, which treats script, style, textarea and title as raw text,
never lets a void element take children, and closes an unclosed element at its parent rather than at the end
of the document. The test suite asserts both of the cases a regular expression gets wrong: a </body>
inside a script string, and an unclosed element in one document.
Duplicates are keyed by what actually makes two elements the same: href for a link, src for a script,
name or property for a meta, and charset on its own. A <style> block has no key, because two style
blocks are two different blocks and dropping one would change the page.
Every document with neither a head nor a body is treated as a fragment and its whole content becomes body, which is what a partial is, and the count of fragments is reported so a mismatched separator is visible rather than mysterious.
Frequently asked questions
What separates one document from the next?
Which title is kept?
Are duplicate inline styles merged?
<style> blocks are kept as two blocks, because they are not necessarily the same thing and
dropping one would change the page. Only linked stylesheets are de-duplicated, by URL.