HTML Code Cleaner

Cleans the markup Word and Google Docs paste: MsoNormal classes, mso- styles, namespaced elements, span soup, empty paragraphs and the document's own fonts.

Enable JavaScript to customise; default output below.

Paste the HTML, not the text: use your editor's code view, or the source pane of whatever you pasted into.

Live preview clean.html
<p>Hello <strong>world</strong>, pasted from Word.</p>

<p>And this one from Google Docs.</p>

<!--
In                              442 characters
Out                             93 characters
Removed                         349 characters (79%)

Comments dropped                1
Namespaced elements dropped     1
Script, style and meta dropped  0
Attributes dropped              7
Wrappers unwrapped              3
Empty elements dropped          1
Tags modernised                 1
Divs turned into paragraphs     1

The Word and Docs classes, the mso- style properties and the namespaced
elements are all the document describing its own formatting. Removing
them is what makes pasted text take on your page's styles instead of
fighting them.

Inline font-family, font-size and line-height were dropped as well.
Those are what make pasted text arrive in Calibri at 11pt, and an inline
style beats anything your stylesheet says, so leaving them in is why the
paragraph never quite matches the page.

Spans and divs left with no attributes are unwrapped rather than kept,
because a wrapper that carries nothing is only there to hold a class
that has been removed.

Classes and styles that were not the processor's were kept. Turn on the
strip option to remove all of them and leave the styling entirely to
your stylesheet.
-->

Output is valid and updates as you type.

Paste out of Word or Google Docs and you get markup that renders and should not be published: class="MsoNormal" on every paragraph, mso- properties in the inline styles, <o:p> and <w:…> elements from XML namespaces nothing else knows about, conditional comments, a <span style="font-family:Calibri"> around every run of text, <p>&nbsp;</p> where somebody pressed return twice, and spans nested four deep for a single bold word.

None of that is a formatting preference. It is the document describing its own fonts and spacing, and because an inline style beats a stylesheet, it wins. That is why pasted text looks almost right and never quite right, and why the usual fix is to retype it.

How to use

  1. Paste the HTML, not the rendered text: your editor’s code view, or the source pane of whatever you pasted into.
  2. Leave “drop inline fonts” on. That is the setting that makes the text take your page’s typography.
  3. Turn on “strip every class, id and style” for body copy. Leave it off if the markup is a layout you built on purpose.

Example

A paragraph from Word and one from Google Docs:

<p>Hello <strong>world</strong>, pasted from Word.</p>
<p>And this one from Google Docs.</p>

That is the whole of it. The input was 442 characters and the output is 93, a 79 percent reduction, and everything removed was the document talking about itself: a conditional comment, a <w:WordDocument> element, four Mso classes, the mso- style properties, font-family:Calibri, font-size:11.0pt, an <o:p> element, an empty paragraph, and three spans that had nothing left on them.

Note what survived: the sentence, the emphasis, and the paragraph break between the two. Word’s <b> became <strong>, and the Google Docs <div> became a <p> rather than being unwrapped, because unwrapping it would have run the two paragraphs together.

Pitfalls

Inline font declarations are the actual problem. Removing MsoNormal and leaving font-family:Calibri;font-size:11pt in place fixes nothing: the text still arrives in the document’s typeface at the document’s size, and your stylesheet cannot override an inline style without !important. That is why dropping them is on by default.

Stripping every class is not always right. For a paragraph of body copy it is exactly right. For markup you wrote, with classes your stylesheet targets, it deletes your work. The option is off by default for that reason.

An empty paragraph is sometimes deliberate spacing. It should not be, and someone used it anyway. If your layout depends on a <p>&nbsp;</p> for spacing, turning off “drop empty elements” keeps it, and the better fix is CSS margin.

Lists from Word are often not lists. Word exports numbered lists as paragraphs with a class, a hanging indent and a literal number or a bullet glyph in the text. The cleaner removes the class and the indent, which leaves paragraphs starting with “1.” Turning those into a real <ol> needs a judgement call about where the list starts and ends, so the tool leaves them alone rather than guessing.

Tracked changes and comments may be in the markup. Word can export both as markup, so check the output before you publish something a client pasted, particularly if the original document had comments in it.

This is not a sanitiser. Event handler attributes and script, style and meta elements are removed, which is a side effect of cleaning rather than a security boundary. Content from an untrusted source has to be sanitised server side by something that was written for it, such as wp_kses.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored. That matters when the paste is a client’s unpublished document.

The markup is parsed rather than pattern-matched, using the same parser as the HTML formatter and the HTML-to-blocks converter on this site. That is what lets it unwrap a span only when the span has nothing left on it, and tell a <div> that wraps blocks from a <div> that is standing in for a paragraph.

The class patterns cover Word (Mso…), Google Docs (OutlineElement, TextRun, NormalTextRun, SCXW…, EOP), LibreOffice (western, font0) and Office Online. Anything that does not match one of those is treated as yours and kept, unless you ask for everything to be stripped.

Input is capped at 100,000 characters. Above that, clean it a section at a time.

Frequently asked questions

Can I paste as plain text instead?
Yes, and it is often the better move: paste with Cmd or Ctrl + Shift + V and add the formatting back yourself. This tool is for when the formatting is worth keeping, or when the paste already happened and the page is full of MsoNormal.
Does WordPress not already clean this?
Partly. The block editor’s paste handling removes a lot of it, and wp_kses strips what is not allowed for the current user. Neither reliably removes the inline font declarations, which is the part you actually see, and content that arrived through an importer or the classic editor may never have been through either.
Why did my <b> become <strong>?
Because <b> is presentational and <strong> carries meaning, and a word processor writes <b> for both. Turn “modernise” off to keep the tags as they are.
What happens to images?
They stay, with their src intact. The width and height attributes are kept unless you strip everything, because for an image those two are worth having.
Will this fix Word’s smart quotes and long dashes?
It leaves them exactly as they are, because they are text rather than markup and they are usually what you want. Curly quotes and an em dash are typographically correct; Word is right about those.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.