Skip to content

HTML entity encoder and decoder

Encode text so it is safe inside HTML, or decode entities back to characters — with the four that actually matter called out.

Minimal is correct for text between tags. Attribute values need the quotes too, and that difference is a real cross-site-scripting hole rather than a style preference.
Runs in your browser — nothing is sent anywhere.

Only three characters actually have to be escaped

Between HTML tags, exactly three: <, > and &. Everything else is a character that renders as itself, and escaping it makes the source harder to read for no benefit.

Inside an attribute value, the quote marks join them. That difference is not a style preference — it is a security boundary. If you write href="…" and the value contains an unescaped double quote, the attribute ends early and whatever comes after it is parsed as markup. That is one of the oldest cross-site-scripting holes there is, and it is why the attribute-safe scope is the default here.

So the scopes are: the three HTML requires, those plus the quotes, those plus everything non-ASCII, or every character that has a name at all. Pick from the top unless you have a reason.

When to escape everything non-ASCII

Rarely, now. It was necessary when a page might be served without a stated character set and a browser would guess. Modern pages declare UTF-8 and accented characters travel perfectly well as themselves.

Where it still helps: email, where some clients still mangle raw UTF-8, and any system that will pass your text through something old and unknown. It makes the output much longer, which is the cost.

Named or numeric

&lt; and &#60; are the same character. The names are easier to read; the numbers work in more places.

The specific case that matters: XML defines only five named entities — lt, gt, amp, quot and apos. Put &nbsp; in an XML document, an SVG or an RSS feed and the parser fails outright, because it has never heard of that name. Numeric entities always work.

Decoding, and the honest limit

The HTML5 specification defines 2,231 named character references. Shipping all of them would add roughly 40 KB to the page for names nobody has typed since 1999, so what is here is the set that appears in real documents: the whole of Latin-1, the curly quotes and dashes a word processor produces, currency, the common maths and arrows, and Greek. About 240 names.

Every numeric form is always decoded, in decimal and hex, so a document using a name that is not in the table still comes through if it uses numbers. And a name that is not recognised is left exactly as written and reported — rather than being silently dropped, which is the failure mode that loses a character without telling you.

An entity without its semicolon

&amp with no semicolon is not quite an entity. A browser decodes some of these anyway — HTML5 keeps a legacy list for the ones people used to write that way — and this does not.

That is a deliberate choice, because the rule is worse than it sounds. Since &not is on that legacy list, HTML5 says &notit; decodes to ¬it; rather than being left alone. Implementing that half-correctly is more dangerous than not implementing it, so the semicolon is required — and anything that looks like a near miss is reported, so you can add the semicolons and try again rather than wondering why nothing happened.

Why your text has &amp;amp; in it

Because it was escaped twice. Each pass turns & into &amp;, so a value escaped on the way into a database and again on the way out arrives with the doubling visible. Decoding once here undoes one layer, and you can see how many layers there are by how many times you have to.

The fix is upstream: escape at the point of output only, never at the point of storage.

It runs in this page

Escaping is what you do to text before putting it somewhere it will be interpreted, which means the text is often a user-submitted value you are investigating, or the contents of a page you are debugging. Nothing here is uploaded.

01

Common questions

Which characters do I actually have to escape?

Between tags, only < > and &. Inside an attribute value, the quote marks as well — an unescaped quote there ends the attribute early and turns the rest into markup.

Named entities or numeric ones?

Names read better in HTML. Use numeric for XML, SVG or RSS: XML knows only five names, so &nbsp; makes an XML parser fail outright.

Why is my text full of &amp;amp;amp;?

It was escaped twice. Each pass turns & into &amp;. Decode once per layer here; the real fix is to escape only at output, never at storage.

Do you support every named entity?

About 240 of the 2,231 in the spec — the ones that appear in real documents. Every numeric form is always decoded, and any name that is not recognised is left as written and reported rather than dropped.

Should I escape accented characters?

Usually not. Modern pages declare UTF-8 and é travels fine as itself. It is worth doing for email, where some clients still mangle raw UTF-8.

Is this the same as URL encoding?

No. HTML entities are for text going into markup; percent-encoding is for text going into a URL. Using one where the other belongs is a common and confusing bug.