HTML Entity Encoder & Decoder
Encode text so markup shows as text, or decode entities back again. Choose how much to encode: markup only, quotes too, or everything above ASCII.
<a href="/plugins?sort=new&page=2">Tom & Jerry's "café" — 50% off</a>
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
Encode text so a browser shows it instead of running it, or decode entities back into the characters they stand for.
How to use
- Paste the text. Encoding turns the characters that mean something to a parser into references; decoding does the reverse.
- Choose how much to encode. Markup only covers the three characters that change structure; add quotes when the text goes inside an attribute value.
- Use “everything above ASCII” when the page is not served as UTF-8, or when a system down the line mangles accents.
- Copy the result. Encoding is safe to repeat: running it over its own output does not double-encode.
Example
<a href="/plugins?sort=new&page=2">Tom & Jerry's "café" — 50% off</a>
With markup and quotes:
<a href="/plugins?sort=new&page=2">Tom & Jerry's "café" — 50% off</a>
The accented é, the em dash and the percent sign are untouched, because none of them change how a parser reads the text.
Pitfalls
- Encoding is not sanitising.
<script>is safe to display, but escaping on the way in to the database and then again on the way out gives you visible&lt;in the page. - The order matters. Ampersands have to be encoded first, or
<becomes&lt;instead of<. That order is what makes a second pass here harmless. 'is HTML5 and XML, but not HTML4. Use'if something very old has to read it.- An attribute value needs its quote character encoded, and nothing else. Encoding the whole attribute is why
class="&quot;x&quot;"shows up in page source. - A bare ampersand in text is technically an error but every browser recovers from it. Decoding here leaves it alone rather than guessing where the entity was meant to end.
- Named references number in the thousands. This decodes the ones that appear in real content; anything else is left visible so you can see what it was.
- Numeric references above the Unicode range, and the surrogate halves, are not characters. They stay as references rather than becoming replacement characters.
- Percent encoding is a different thing.
%20is for URLs, and belongs to the URL encoder.
Compatibility
The five markup references are defined by both HTML and XML, so they work everywhere. Numeric references, decimal or hexadecimal, are equally universal. The named references decoded here are the HTML5 set that appears in real content. Everything runs in your browser, with no upload and no DOM parsing, so the behaviour is the same in every browser from 2017 onwards.