HTML Code Cleaner
Cleans the markup Word and Google Docs paste: MsoNormal classes, mso- styles, namespaced elements, span soup, empty paragraphs and the document's own fonts.
<p>Hello <strong>world</strong>, pasted from Word.</p>
<p>And this one from Google Docs.</p>
<!--
In 442 characters
Out 93 characters
Removed 349 characters (79%)
Comments dropped 1
Namespaced elements dropped 1
Script, style and meta dropped 0
Attributes dropped 7
Wrappers unwrapped 3
Empty elements dropped 1
Tags modernised 1
Divs turned into paragraphs 1
The Word and Docs classes, the mso- style properties and the namespaced
elements are all the document describing its own formatting. Removing
them is what makes pasted text take on your page's styles instead of
fighting them.
Inline font-family, font-size and line-height were dropped as well.
Those are what make pasted text arrive in Calibri at 11pt, and an inline
style beats anything your stylesheet says, so leaving them in is why the
paragraph never quite matches the page.
Spans and divs left with no attributes are unwrapped rather than kept,
because a wrapper that carries nothing is only there to hold a class
that has been removed.
Classes and styles that were not the processor's were kept. Turn on the
strip option to remove all of them and leave the styling entirely to
your stylesheet.
-->
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Paste out of Word or Google Docs and you get markup that renders and should not be
published: class="MsoNormal" on every paragraph, mso- properties in the inline
styles, <o:p> and <w:…> elements from XML namespaces nothing else knows about,
conditional comments, a <span style="font-family:Calibri"> around every run of
text, <p> </p> where somebody pressed return twice, and spans nested four deep
for a single bold word.
None of that is a formatting preference. It is the document describing its own fonts and spacing, and because an inline style beats a stylesheet, it wins. That is why pasted text looks almost right and never quite right, and why the usual fix is to retype it.
How to use
- Paste the HTML, not the rendered text: your editor’s code view, or the source pane of whatever you pasted into.
- Leave “drop inline fonts” on. That is the setting that makes the text take your page’s typography.
- Turn on “strip every class, id and style” for body copy. Leave it off if the markup is a layout you built on purpose.
Example
A paragraph from Word and one from Google Docs:
<p>Hello <strong>world</strong>, pasted from Word.</p>
<p>And this one from Google Docs.</p>
That is the whole of it. The input was 442 characters and the output is 93, a 79
percent reduction, and everything removed was the document talking about itself: a
conditional comment, a <w:WordDocument> element, four Mso classes, the mso-
style properties, font-family:Calibri, font-size:11.0pt, an <o:p> element, an
empty paragraph, and three spans that had nothing left on them.
Note what survived: the sentence, the emphasis, and the paragraph break between the
two. Word’s <b> became <strong>, and the Google Docs <div> became a <p>
rather than being unwrapped, because unwrapping it would have run the two paragraphs
together.
Pitfalls
Inline font declarations are the actual problem. Removing MsoNormal and leaving
font-family:Calibri;font-size:11pt in place fixes nothing: the text still arrives in
the document’s typeface at the document’s size, and your stylesheet cannot override an
inline style without !important. That is why dropping them is on by default.
Stripping every class is not always right. For a paragraph of body copy it is exactly right. For markup you wrote, with classes your stylesheet targets, it deletes your work. The option is off by default for that reason.
An empty paragraph is sometimes deliberate spacing. It should not be, and someone
used it anyway. If your layout depends on a <p> </p> for spacing, turning off
“drop empty elements” keeps it, and the better fix is CSS margin.
Lists from Word are often not lists. Word exports numbered lists as paragraphs
with a class, a hanging indent and a literal number or a bullet glyph in the text. The
cleaner removes the class and the indent, which leaves paragraphs starting with “1.”
Turning those into a real <ol> needs a judgement call about where the list starts and
ends, so the tool leaves them alone rather than guessing.
Tracked changes and comments may be in the markup. Word can export both as markup, so check the output before you publish something a client pasted, particularly if the original document had comments in it.
This is not a sanitiser. Event handler attributes and script, style and meta
elements are removed, which is a side effect of cleaning rather than a security
boundary. Content from an untrusted source has to be sanitised server side by something
that was written for it, such as wp_kses.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored. That matters when the paste is a client’s unpublished document.
The markup is parsed rather than pattern-matched, using the same parser as the HTML
formatter and the HTML-to-blocks converter on this site. That is what lets it unwrap a
span only when the span has nothing left on it, and tell a <div> that wraps blocks
from a <div> that is standing in for a paragraph.
The class patterns cover Word (Mso…), Google Docs (OutlineElement, TextRun,
NormalTextRun, SCXW…, EOP), LibreOffice (western, font0) and Office Online.
Anything that does not match one of those is treated as yours and kept, unless you ask
for everything to be stripped.
Input is capped at 100,000 characters. Above that, clean it a section at a time.
Frequently asked questions
Can I paste as plain text instead?
MsoNormal.Does WordPress not already clean this?
wp_kses strips
what is not allowed for the current user. Neither reliably removes the inline font
declarations, which is the part you actually see, and content that arrived through an
importer or the classic editor may never have been through either.Why did my <b> become <strong>?
<b> is presentational and <strong> carries meaning, and a word processor
writes <b> for both. Turn “modernise” off to keep the tags as they are.What happens to images?
src intact. The width and height attributes are kept unless
you strip everything, because for an image those two are worth having.