URL Encoding Explained: Percent-Encoding, encodeURIComponent and +
Drafted with AI assistance. Every command and code example was run and its output checked before publication. How guides are made
URL encoding, properly called percent-encoding, replaces each byte that may not appear literally in a part of a URL with % followed by two hexadecimal digits, after first converting the text to UTF-8. A space becomes %20, & becomes %26 and é becomes %C3%A9. In JavaScript, encode each value with encodeURIComponent or let URLSearchParams build the query string; reserve encodeURI for tidying a complete address you already trust. The one place a space may be written as + is form-encoded data such as an HTML form submission.
What is percent-encoding?
RFC 3986, the generic URI syntax, defines a percent-encoded octet as % plus two hex digits giving the byte's value (section 2.1). The scheme works on bytes, not characters, so text has to be turned into bytes first. Modern specifications and every function in this guide use UTF-8, which is why one character can become several triplets:
encodeURIComponent("é"); // "%C3%A9" (2 UTF-8 bytes)
encodeURIComponent("€"); // "%E2%82%AC" (3 bytes)
encodeURIComponent("😀"); // "%F0%9F%98%80" (4 bytes)
Hex digits are case-insensitive: %c3%a9 and %C3%A9 mean the same thing, and the RFC asks producers to use upper case. The percent sign itself is data only when written as %25.
Which characters must be URL encoded?
RFC 3986 splits the printable ASCII characters into two groups (sections 2.2 and 2.3):
- Unreserved:
A–Z a–z 0–9 - . _ ~. These never need encoding, and the RFC says producers should not encode them. - Reserved: the general delimiters
: / ? # [ ] @and the sub-delimiters! $ & ' ( ) * + , ; =. They have a job in URL syntax, so when one of them is part of your data it must be encoded, otherwise it is read as structure.
Everything else, including the space, ", <, >, %, control characters and all non-ASCII text, is not allowed literally and is always encoded. The practical rule follows from the reserved set: a / inside a file name, an & or = inside a query value and a # anywhere in the data must be encoded, or the server sees an extra path segment, an extra parameter or the start of a fragment.
encodeURI vs encodeURIComponent vs URLSearchParams
JavaScript has three built-in ways to percent-encode, and they leave different characters alone. The table lists exactly which printable ASCII characters each one passes through unchanged, as checked in Node.js 20; every other character is encoded.
| Function | Left unencoded (besides letters and digits) | Space becomes | Use it for |
|---|---|---|---|
encodeURIComponent | - . _ ~ ! ' ( ) * | %20 | One path segment, one query name or value |
encodeURI | - . _ ~ ! ' ( ) * ; , / ? : @ & = + $ # | %20 | A whole URL that is already correctly structured |
URLSearchParams | - . _ * | + | Building or reading a query string |
| RFC 3986 unreserved set | - . _ ~ | %20 | Strict encoding, e.g. OAuth 1.0 signatures |
Note what encodeURI does not protect: because it keeps &, =, ? and #, a value run through it can still break the query. Both functions encode [ and ] as %5B and %5D, even though they are reserved characters.
const q = "café & crème";
encodeURI("https://example.com/search?q=" + q);
// "https://example.com/search?q=caf%C3%A9%20&%20cr%C3%A8me" (& still splits the query)
"https://example.com/search?q=" + encodeURIComponent(q);
// "https://example.com/search?q=caf%C3%A9%20%26%20cr%C3%A8me"
The rule of thumb: encode each piece of data on its own with encodeURIComponent, then join the pieces with literal /, ?, & and =. The behaviour of both functions is fixed by the ECMAScript specification, so every browser and Node.js gives the same output. You can compare the two modes side by side in the URL Encoder / Decoder, which calls exactly these functions.
Building URLs with URL and URLSearchParams
For query strings it is usually simpler to let the platform do the joining. URL and URLSearchParams follow the WHATWG URL Standard, and they encode values as they are set:
const url = new URL("https://example.com/search");
url.searchParams.set("q", "café & crème");
url.searchParams.set("page", "2");
url.href;
// "https://example.com/search?q=caf%C3%A9+%26+cr%C3%A8me&page=2"
new URLSearchParams("q=caf%C3%A9+%26+cr%C3%A8me").get("q");
// "café & crème"
A whole URL passed as a parameter, such as a redirect value, is encoded safely too: its :, /, ? and & all become escapes, so it travels as one value.
Should a space be %20 or +?
Both are correct, in different places. RFC 3986 knows only %20; a + in a path is a literal plus sign. The + for a space comes from the application/x-www-form-urlencoded format, which browsers use when they submit an HTML form and which URLSearchParams implements. That format encodes a real plus sign as %2B and decodes + back to a space.
Problems appear when the two conventions are mixed. decodeURIComponent does not know about the form rule, so it leaves + alone:
decodeURIComponent("a+b%20c"); // "a+b c"
new URLSearchParams("q=a+b%20c").get("q"); // "a b c"
If you read a query string with decodeURIComponent, spaces submitted by a form stay as plus signs. If you build one by hand and put a literal + in a value, a form-aware server reads it as a space; C++ arrives as C . Encode values with encodeURIComponent (which turns + into %2B) or URLSearchParams, and decode query strings with URLSearchParams.
What is double encoding, and how do I fix %2520?
Double encoding happens when text that is already percent-encoded is encoded again. The % of each escape becomes %25, so %20 turns into %2520:
encodeURIComponent(encodeURIComponent("a b")); // "a%2520b"
encodeURI("https://example.com/a%20b"); // "https://example.com/a%2520b"
RFC 3986 section 2.4 is explicit: implementations must not percent-encode or decode the same string more than once. The usual cause is two layers each encoding "to be safe", for example a helper that encodes a parameter and an HTTP client that encodes the whole URL again. Decide which layer owns encoding, keep values raw until that point, and encode exactly once. To repair a value that is already double-encoded, decode it twice. The reverse mistake, decoding twice, is a security concern: a value that has been checked once, then decoded again, can smuggle characters such as %2F or %3C past the check.
Why does decodeURIComponent throw "URI malformed"?
decodeURIComponent and decodeURI throw a URIError when a % is not followed by two hex digits, or when the bytes do not form valid UTF-8. The message in V8 (Chrome and Node.js) is URI malformed; other engines word it differently.
decodeURIComponent("100%"); // URIError: URI malformed (bare percent sign)
decodeURIComponent("%E9"); // URIError: URI malformed (Latin-1 é, not UTF-8)
encodeURIComponent("\uD800"); // URIError: URI malformed (lone surrogate)
The first case is usually text that was never encoded; the second is a URL produced by an old system that encoded Latin-1 instead of UTF-8. Wrap decoding of untrusted input in try/catch and treat a failure as bad input rather than letting it crash a request handler:
function safeDecode(s) {
try { return decodeURIComponent(s); } catch { return null; }
}
Other languages are more forgiving, which can hide the problem. Python's urllib.parse.unquote("%E9") returns the replacement character �, and unquote("100%") returns 100% unchanged.
decodeURI also behaves differently from decodeURIComponent: it leaves escapes that stand for reserved characters encoded, so decodeURI("%3F%20x") gives %3F x, while decodeURIComponent gives ? x.
How do I URL encode in Python and PHP?
Python's urllib.parse has two encoders. quote follows RFC 3986 but treats / as safe by default, because it is meant for paths; pass safe="" to encode a single segment or value. quote_plus uses the form convention, and urlencode builds a whole query string from a dictionary with quote_plus:
from urllib.parse import quote, quote_plus, urlencode, unquote_plus
quote("café & crème/x") # 'caf%C3%A9%20%26%20cr%C3%A8me/x'
quote("café & crème/x", safe="") # 'caf%C3%A9%20%26%20cr%C3%A8me%2Fx'
quote_plus("café & crème/x") # 'caf%C3%A9+%26+cr%C3%A8me%2Fx'
urlencode({"q": "café & crème", "tag": "a+b"})
# 'q=caf%C3%A9+%26+cr%C3%A8me&tag=a%2Bb'
unquote_plus("a+b%20c") # 'a b c'
With safe="", quote leaves exactly the RFC 3986 unreserved characters alone, which makes it stricter than encodeURIComponent.
PHP mirrors the same split. rawurlencode follows RFC 3986 (space as %20, ~ kept), while urlencode uses the form convention (space as +, and ~ encoded as %7E):
rawurlencode("café & crème~"); // caf%C3%A9%20%26%20cr%C3%A8me~
urlencode("café & crème~"); // caf%C3%A9+%26+cr%C3%A8me%7E
urldecode("a+b%20c"); // a b c
rawurldecode("a+b%20c"); // a+b c
PHP works on the bytes of the string, so these outputs assume the source file and data are UTF-8. http_build_query builds a query string from an array in the form style, like Python's urlencode.
How are non-ASCII domain names encoded?
Not with percent-encoding. Host names that contain non-ASCII characters (internationalised domain names) are converted label by label to an ASCII form with the xn-- prefix, using the Punycode algorithm from RFC 3492 under the IDNA rules of RFC 5891. DNS only ever sees the ASCII form. The path, query and fragment of the same URL are still percent-encoded as UTF-8. The WHATWG URL parser does both for you:
new URL("https://münchen.de/straße?q=a b&x=ü").href;
// "https://xn--mnchen-3ya.de/stra%C3%9Fe?q=a%20b&x=%C3%BC"
This is also why encodeURI is the wrong tool for a URL with a non-ASCII host: it percent-encodes the host name (m%C3%BCnchen.de), which is not the form DNS uses. A WHATWG URL parser repairs that to xn--mnchen-3ya.de, but other software may not.
Is URL encoding the same as Base64 or HTML escaping?
No. Each encoding protects text for a different parser. Percent-encoding makes data safe inside a URL. Base64 turns arbitrary bytes into text, and its standard alphabet includes +, / and =, so a Base64 value placed in a query string still needs percent-encoding, or the Base64URL variant. HTML escaping protects characters that are special to the HTML parser; a URL written into an href attribute needs percent-encoding for its parts and then HTML escaping for the attribute, which Escaping HTML correctly covers in detail.
Quick reference
- Encode each value separately with
encodeURIComponent, then join with literal/ ? & =. - Prefer
URLandURLSearchParamsfor query strings; they encode onsetand decode onget. - Use
encodeURIonly on a complete, trusted URL; it does not protect&,=,?or#inside values. %20is always a space;+is a space only in form-encoded data. A literal plus in a value must be%2B.- Encode once and decode once.
%2520means something encoded twice. - Catch
URIErrorwhen decoding untrusted input. - Python:
quote(s, safe="")for a value,urlencodefor a query. PHP:rawurlencodefor paths,http_build_queryfor queries. - Non-ASCII host names use Punycode (
xn--), not percent-encoding.
Frequently asked questions
Why does encodeURIComponent not encode ! ' ( ) and *?
The ECMAScript definition treats them as safe, although RFC 3986 lists them as reserved sub-delimiters. Most servers accept them literally. Where strict RFC 3986 encoding is required, such as OAuth 1.0 signature base strings, replace them afterwards: s.replace(/[!'()*]/g, c => "%" + c.charCodeAt(0).toString(16).toUpperCase()).
Do I need to encode a hyphen, dot, underscore or tilde?
No. They are unreserved, and RFC 3986 says producers should not encode them. %7E and ~ are equivalent, although URLSearchParams and PHP's urlencode still write %7E.
Are %c3%a9 and %C3%A9 the same?
Yes. Hex digits in an escape are case-insensitive, so the two URLs are equivalent. Upper case is recommended when producing URLs, and it is what JavaScript, Python and PHP output.
Can I put a URL inside another URL's query string?
Yes, if the inner URL is encoded as a value. With encodeURIComponent or URLSearchParams its :, /, ? and & become escapes, so the outer URL treats it as one parameter. If you only use encodeURI, the inner & splits it into several parameters.
Is URL encoding a security measure?
It prevents data from being misread as URL structure, which matters for correctness and for injection into URLs. It does not make a URL safe to place in a page: encodeURI("javascript:alert(1)") returns the input unchanged, so validate the scheme of untrusted URLs as well.