Escaping HTML Correctly: Entities, Contexts and XSS

By , founder of Softaware Commerce · Published · Updated

Drafted with AI assistance. Every command and code example was run and its output checked before publication. How guides are made

To escape text for HTML, replace & with &amp;, < with &lt;, > with &gt;, " with &quot; and ' with &#39;, ampersand first. That is enough for text between tags and for attribute values in quotes, but not for anywhere else: URLs in href, inline <script> blocks, event handler attributes and CSS each have their own rules. Cross-site scripting (XSS) happens when untrusted text reaches one of those places without the encoding that place needs, which is why escaping must match the output context.

Which characters need escaping in HTML?

Five characters can change how the HTML parser reads text:

CharacterNamed referenceNumeric referenceWhy it matters
&&amp;&#38;Starts a character reference
<&lt;&#60;Starts a tag or comment
>&gt;&#62;Ends a tag; escaped for consistency
"&quot;&#34;Ends a double-quoted attribute value
'&apos;&#39;Ends a single-quoted attribute value

The order matters: escape & first, or you will turn the & of every reference you just wrote into &amp;. A minimal JavaScript escaper, the same replacement chain the HTML Entity Encoder runs when you click Encode, looks like this:

function escapeHtml(s) {
  return s.replace(/&/g, "&amp;")
          .replace(/</g, "&lt;")
          .replace(/>/g, "&gt;")
          .replace(/"/g, "&quot;")
          .replace(/'/g, "&#39;");
}

Accented letters, symbols and emoji do not need escaping on a page served as UTF-8; they are not special to the parser.

Named or numeric entities?

Both forms produce the same character. &lt;, &#60; and &#x3C; all decode to <. The HTML Standard's list of named character references is long, but only a handful matter for escaping. Numeric references work for any Unicode code point and need no lookup table. &apos; is valid in HTML today but was not defined in HTML 4, which is why many escapers, including Python's and PHP's, emit a numeric form for the apostrophe instead:

import html
html.escape("<a href=\"x\">Tom & Jerry's</a>")
# '&lt;a href=&quot;x&quot;&gt;Tom &amp; Jerry&#x27;s&lt;/a&gt;'
<?php
echo htmlspecialchars("<a href=\"x\">Tom & Jerry's</a>");
// &lt;a href=&quot;x&quot;&gt;Tom &amp; Jerry&#039;s&lt;/a&gt;  (PHP 8.1+ default flags)

Why does escaping depend on the context?

The browser does not run one parser over a page; it runs several. The HTML parser reads the markup, then hands attribute values that are URLs to the URL parser, the contents of <script> and event handlers to the JavaScript engine, and style values to the CSS parser. Each has its own special characters. The OWASP Cross Site Scripting Prevention Cheat Sheet organises its rules by these output contexts, and notes that using the wrong encoding method may introduce weaknesses. The examples below were each run in Chromium; escapeHtml is the function above.

Where the value goesIs HTML escaping enough?What to do
Between tags: <p>VALUE</p>YesEscape the five characters, or set textContent
Quoted attribute: title="VALUE"YesEscape, and always quote the attribute
Unquoted attribute: title=VALUENoAdd quotes
URL attribute: href="VALUE"NoCheck the scheme, percent-encode parts, then escape
Event handler: onclick="f('VALUE')"NoUse addEventListener and pass data
Inside <script>NoJSON-encode and escape <, or use a data attribute
CSS: style="color: VALUE"NoAllowlist the value

Element content

Unescaped, <img src=x onerror=alert(1)> inserted between tags becomes an image whose error handler runs. Escaped, it is displayed as text. This is the case most escaping functions are built for.

Attribute values: quoted and unquoted

In a quoted attribute, the dangerous character is the quote. Without escaping, the value " autofocus onfocus="alert(1) closes value=" early and adds two attributes. Escaped, the " becomes &quot; and the whole string stays inside value.

Unquoted attributes end at whitespace, and none of the five escapes touch spaces. So <input value=VALUE> with x autofocus onfocus=alert(1) still produces three attributes after escaping. OWASP's rule is simply to always quote attribute values.

URLs in href and src

HTML escaping does nothing to javascript:alert(1): it contains none of the five characters, so <a href="javascript:alert(1)"> runs the script when clicked. Entity-encoding the scheme does not hide it either, because the parser decodes references in attributes before the URL is used: href="javascript&#58;alert(1)" also runs. For a user-supplied URL, parse it and allow only the schemes you expect:

function safeUrl(input) {
  try {
    const url = new URL(input, location.href);
    return ["http:", "https:", "mailto:"].includes(url.protocol) ? url.href : "about:blank";
  } catch {
    return "about:blank";
  }
}
safeUrl(" JaVaScRiPt:alert(1)");  // "about:blank"

Parsing with URL matters, because browsers ignore leading spaces, letter case and embedded tabs in the scheme; a hand-written startsWith("javascript:") check misses those. When you build a URL from parts, percent-encode each part (see URL encoding explained and the URL Encoder), then HTML-escape the finished URL for the attribute.

Event handler attributes

An onclick value is HTML first and JavaScript second, and the HTML parser decodes references before the script runs. With onclick="greet('VALUE')" and the value ');alert(2);//, HTML escaping turns the apostrophe into &#39;, the parser turns it back into ', and the injected call runs. OWASP's advice is to avoid this context: attach handlers with addEventListener and pass the value as data.

Inline scripts and JSON in a script tag

Inside <script>, HTML references are not decoded at all, so "Tom &amp; Jerry" stays as those exact characters. Yet the HTML parser still looks for </script to find the end of the element, before JavaScript sees anything. JSON.stringify does not escape < or /, so this common pattern is injectable:

// server-side template
`<script>var user = ${JSON.stringify(user)};</script>`
// user.name = "</script><script>alert(1)</script>"  → the alert runs

The browser closes the first script in the middle of the string and runs the second one as a new script. A related trap is <!--: per the HTML Standard's restrictions for contents of script elements, the sequence <!--<script> in a script puts the parser into a state where the next </script> no longer ends the element, so the rest of the page is swallowed into the script. The standard's own advice is to write those sequences as \x3C!--, \x3Cscript and \x3C/script. For JSON, escaping every <, > and & as a Unicode escape covers all three cases and still parses to the same value:

function jsonForScript(value) {
  return JSON.stringify(value)
    .replace(/</g, "\\u003c")
    .replace(/>/g, "\\u003e")
    .replace(/&/g, "\\u0026");
}

Writing <\/script>, as the String Escape tool's page suggests, fixes the closing-tag case but not the <!-- case. Often the cleanest option is not to generate script at all: put the JSON in <script type="application/json" id="data"> (still escaped as above) or in a data- attribute, and read it with JSON.parse(el.textContent).

CSS

OWASP's guidance is that untrusted data should only ever go into a CSS property value, never a selector or a whole declaration, and that properties that take URLs need URL validation too. In practice, check the value against an allowlist (a set of colour names, or a pattern such as /^#[0-9a-f]{6}$/i) and set it with el.style.color = value rather than by building a style string.

textContent vs innerHTML: which should I use?

When writing text with the DOM, you usually do not need to escape anything. textContent creates a text node, so the browser never parses the value as HTML:

const p = document.createElement("p");
p.textContent = '<img src=x onerror=alert(1)>';   // shown as text, nothing runs
p.innerHTML = '<img src=x onerror=alert(1)>';     // creates an img; the alert runs

Do not escape before assigning to textContent; you will see the references on screen. innerHTML does not execute <script> elements it inserts, but event handler attributes such as onerror do run, so it is never safe for untrusted strings. The same goes for outerHTML, insertAdjacentHTML and document.write. Safe alternatives are textContent, value, setAttribute for non-URL, non-handler attributes, and building elements with createElement.

Do template engines escape automatically?

Most do, for the HTML text and quoted attribute contexts, and you should rely on that rather than escaping by hand. Two examples, run with Handlebars 4.7 and React 18.3:

  • Handlebars escapes {{value}} (including = and backticks) and outputs {{{value}}} raw.
  • React escapes strings rendered as children or props. dangerouslySetInnerHTML inserts raw HTML, and an href of javascript:alert(1) is still rendered by React 18, with only a console warning.

The gaps are the same everywhere: raw-output syntax, URL attributes, inline scripts and styles. Search a code base for the raw-output forms (such as Handlebars' {{{, React's dangerouslySetInnerHTML, Jinja's and Django's |safe filter and Vue's v-html) and treat each as a review point.

Should I sanitise or escape HTML?

Escape when the value is text, which is almost always. Sanitise only when users are meant to supply HTML, such as rich-text comments or content from a WYSIWYG editor, where escaping would display the tags instead of applying them. A sanitiser parses the HTML and removes anything not on an allowlist. OWASP recommends DOMPurify; with version 3.4 in Chromium:

DOMPurify.sanitize('<p onclick="alert(1)">Hi <b>there</b><img src=x onerror=alert(2)>' +
  '<a href="javascript:alert(3)">l</a><script>alert(4)</script></p>');
// '<p>Hi <b>there</b><img src="x"><a>l</a></p>'

Sanitise as close to output as possible, do not modify the result afterwards, and keep the library up to date. Never write your own sanitiser with regular expressions; HTML parsing has too many edge cases.

What is double escaping, and how do I avoid it?

Double escaping shows up as &amp;, &amp;lt; or &amp;quot; on screen. It happens when text is escaped twice, typically once when it is saved and again by the template: escapeHtml(escapeHtml("Tom & Jerry")) gives Tom &amp;amp; Jerry, which displays as Tom &amp; Jerry. The fix is a rule, not a function: store raw text, and escape once, at output, for the context you are writing into. Some escapers can skip existing references (PHP's htmlspecialchars with double_encode: false), but that hides the bug and cannot represent text that genuinely contains &amp;. To clean up data that is already escaped, decode it once with the HTML Entity Encoder's Decode button.

Checklist

  • Store raw text; escape at output, once, for the exact context.
  • Let the template engine auto-escape; review every raw-output construct.
  • Always quote attribute values.
  • Validate user URLs with new URL() and a scheme allowlist before putting them in href or src.
  • Never build event handler strings from data; use addEventListener.
  • JSON in a page: escape <, > and & as \u003c, \u003e, \u0026, or use a JSON script block or data attribute.
  • CSS values from users: allowlist only.
  • In the DOM, use textContent, not innerHTML.
  • User-authored HTML: sanitise with DOMPurify, never with regex.

Frequently asked questions

Do I need to escape the greater-than sign?

In element content and quoted attributes, > alone cannot start markup, so leaving it is not an injection by itself. Escaping it costs nothing and keeps the rule simple, so common escapers, including Python's html.escape and PHP's htmlspecialchars, always do.

Is escaping the same as encoding?

In this context the words are used interchangeably: OWASP calls it output encoding. What matters is the target. HTML escaping uses character references, URL encoding uses %XX and JavaScript strings use backslash escapes; the String Escape tool produces the last kind.

Is it safer to escape input when it arrives?

No. At input time you do not know where the text will be written, and HTML escaping is wrong for SQL, JSON, URLs or a plain-text email. Validate input for what it should be, store it raw, and encode at output.

Does a Content Security Policy replace escaping?

No. A strict CSP limits what injected markup can do, for example by blocking inline scripts, and is a valuable second layer. It does not stop every injection, such as markup that changes the page's content or links, so escape correctly as well.

Why does my page show &amp;lt; instead of <?

The text was escaped twice. Find the two places that escape it, usually the code that saves the data and the template that prints it, and remove the first.

Tools for this guide