HTML to Text Converter (Tag Remover)

Strips tags by parsing, not by regex: paragraphs stay paragraphs, list items keep markers, table rows stay on one line, script contents stay out.

Enable JavaScript to customise; default output below.

A fragment or a whole document. Head, script, style and svg contents are dropped.

Link addresses

Characters. Zero leaves the lines as long as they are; 72 or 76 is the convention for plain-text email.

Live preview text.txt
Release notes

Version 2.4 is out. Read the changelog[1] for the details.

- Faster sync
- Fixed the import bug

Plan	Price
Pro	$99

Done.
Really.

[1] https://swiftplugins.pro/changelog

----------------------------------------
HTML in     348 characters
Text out    186 characters
Words       28
Links kept  1

The tags were parsed rather than stripped, so a paragraph is a blank
line, a list item keeps its marker and a table row stays on one line.
Script and style contents are dropped rather than read, which is what a
regular expression over the source cannot do.

Output is valid and updates as you type.

Stripping tags with a regular expression gets the words right and loses the shape. Paragraphs run into each other, list items lose their markers, table cells lose their separation, and the contents of a <script> block end up in the middle of a sentence.

This parses the HTML instead. A block element becomes a blank line, a list item becomes a marker, a table row stays on one line with tabs between the cells, a <br> becomes a newline, and script, style and svg contents are dropped rather than read.

How to use

  1. Paste the HTML. A fragment or a whole document, either way.
  2. Choose what happens to link addresses: footnoted, inline, or dropped.
  3. Set a wrap width if this is going into a plain-text email. 72 is the convention.

Example

Release notes

Version 2.4 is out. Read the changelog[1] for the details.

- Faster sync
- Fixed the import[2] bug

Plan	Price
Pro	$99

Done.
Really.

[1] https://swiftplugins.pro/changelog
[2] https://swiftplugins.pro/docs

The footnote form is what mail clients do, and it is the right default: the sentence still reads as a sentence, and the addresses are all in one place at the end. A repeated address reuses its number.

Pitfalls

Regular expressions cannot do this. replace(/<[^>]*>/g, '') breaks on an attribute containing a >, reads the contents of a script tag as prose, and produces one long paragraph. The reason to parse is not purity, it is that the output is different and better.

Whitespace is collapsed, except in pre. That is what a browser does: runs of spaces and newlines in the source are one space. Inside <pre> the text is kept as written, which is what you want for code.

Entities are the parser’s job. &amp; comes out as & and &nbsp; as a space, because that is what the reader saw. If you need the source form, you wanted a different tool.

A <div> is a paragraph break here. Plenty of HTML uses divs where it means paragraphs, so treating them as blocks produces better text more often than treating them as inline. If your markup uses divs as inline wrappers, expect extra blank lines.

Tabs between table cells are not a table. They line up in a monospaced viewer and not in an email client. If the table matters, the honest options are to keep it as HTML or to rewrite it as a list.

Hidden content is still content. An element with display: none in a stylesheet is invisible to a reader and visible to this parser, because the parser does not run CSS. Preview text hidden at the top of an email is the classic case.

Compatibility

Everything runs in the browser: nothing is uploaded and nothing is stored. That matters for pasting a customer email or a page of a client’s site.

The parser is the same one the HTML formatter and the HTML-to-blocks converter on this site use: a small tokeniser that handles void elements, raw-text elements, implicit close tags and unquoted attributes. It does not build a DOM, so it does not need a browser, and the same code runs at build time to produce the default output on this page.

Input is capped at 100,000 characters. Above that, split the document: the parser copes, but a browser holding both the input and the output in a textarea starts to feel it.

Frequently asked questions

Will it keep my line breaks?
Yes, where the HTML has them: <br> becomes a newline, block elements become blank lines, and <pre> keeps its own spacing. Newlines in the HTML source between tags are whitespace and get collapsed, as they do in a browser.
What happens to images?
They produce nothing. An image has no text, and inserting its alt text mid-sentence usually reads worse than leaving it out. If the alt text matters, add it to the HTML as a caption.
Can I use this for plain-text email?
That is what the wrap and the footnote options are for. Set the wrap to 72 and keep the footnotes, which is the convention every mail client expects.
Does it remove HTML comments?
Yes, along with doctypes and the head. Comments are not content.
Is this the same as WordPress’s wp_strip_all_tags?
No. wp_strip_all_tags removes the tags and leaves the text jammed together, and it also removes script and style contents, which is the part most regex approaches get wrong. This adds the structure back as layout, which is the part wp_strip_all_tags does not attempt.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.