HTML to Text Converter (Tag Remover)
Strips tags by parsing, not by regex: paragraphs stay paragraphs, list items keep markers, table rows stay on one line, script contents stay out.
Release notes
Version 2.4 is out. Read the changelog[1] for the details.
- Faster sync
- Fixed the import bug
Plan Price
Pro $99
Done.
Really.
[1] https://swiftplugins.pro/changelog
----------------------------------------
HTML in 348 characters
Text out 186 characters
Words 28
Links kept 1
The tags were parsed rather than stripped, so a paragraph is a blank
line, a list item keeps its marker and a table row stays on one line.
Script and style contents are dropped rather than read, which is what a
regular expression over the source cannot do.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Stripping tags with a regular expression gets the words right and loses the shape.
Paragraphs run into each other, list items lose their markers, table cells lose
their separation, and the contents of a <script> block end up in the middle of a
sentence.
This parses the HTML instead. A block element becomes a blank line, a list item
becomes a marker, a table row stays on one line with tabs between the cells, a <br>
becomes a newline, and script, style and svg contents are dropped rather than read.
How to use
- Paste the HTML. A fragment or a whole document, either way.
- Choose what happens to link addresses: footnoted, inline, or dropped.
- Set a wrap width if this is going into a plain-text email. 72 is the convention.
Example
Release notes
Version 2.4 is out. Read the changelog[1] for the details.
- Faster sync
- Fixed the import[2] bug
Plan Price
Pro $99
Done.
Really.
[1] https://swiftplugins.pro/changelog
[2] https://swiftplugins.pro/docs
The footnote form is what mail clients do, and it is the right default: the sentence still reads as a sentence, and the addresses are all in one place at the end. A repeated address reuses its number.
Pitfalls
Regular expressions cannot do this. replace(/<[^>]*>/g, '') breaks on an
attribute containing a >, reads the contents of a script tag as prose, and produces
one long paragraph. The reason to parse is not purity, it is that the output is
different and better.
Whitespace is collapsed, except in pre. That is what a browser does: runs of
spaces and newlines in the source are one space. Inside <pre> the text is kept as
written, which is what you want for code.
Entities are the parser’s job. & comes out as & and as a space,
because that is what the reader saw. If you need the source form, you wanted a
different tool.
A <div> is a paragraph break here. Plenty of HTML uses divs where it means
paragraphs, so treating them as blocks produces better text more often than treating
them as inline. If your markup uses divs as inline wrappers, expect extra blank lines.
Tabs between table cells are not a table. They line up in a monospaced viewer and not in an email client. If the table matters, the honest options are to keep it as HTML or to rewrite it as a list.
Hidden content is still content. An element with display: none in a stylesheet
is invisible to a reader and visible to this parser, because the parser does not run
CSS. Preview text hidden at the top of an email is the classic case.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored. That matters for pasting a customer email or a page of a client’s site.
The parser is the same one the HTML formatter and the HTML-to-blocks converter on this site use: a small tokeniser that handles void elements, raw-text elements, implicit close tags and unquoted attributes. It does not build a DOM, so it does not need a browser, and the same code runs at build time to produce the default output on this page.
Input is capped at 100,000 characters. Above that, split the document: the parser copes, but a browser holding both the input and the output in a textarea starts to feel it.
Frequently asked questions
Will it keep my line breaks?
<br> becomes a newline, block elements become blank
lines, and <pre> keeps its own spacing. Newlines in the HTML source between tags are
whitespace and get collapsed, as they do in a browser.What happens to images?
Can I use this for plain-text email?
Does it remove HTML comments?
Is this the same as WordPress’s wp_strip_all_tags?
wp_strip_all_tags removes the tags and leaves the text jammed together, and it
also removes script and style contents, which is the part most regex approaches get
wrong. This adds the structure back as layout, which is the part wp_strip_all_tags
does not attempt.