Base64 Encoding Explained: Padding, Base64URL and When to Use It

By the CodeBeautify team at Softaware Commerce Ltd · Published

Base64 turns arbitrary bytes into text built from 64 safe ASCII characters by taking 3 bytes (24 bits) at a time and writing them as 4 characters of 6 bits each, so the output is about a third larger than the input. Padding with = fills out the last group, and Base64URL is a variant that swaps + and / for - and _ so the result is safe in URLs and file names. Use it to carry binary data through text-only channels such as JSON, email or data URIs; it is an encoding, not encryption, and anyone can decode it.

How Base64 works: 3 bytes become 4 characters

Base64 is defined in RFC 4648. The encoder reads the input as a stream of bits, cuts it into 6-bit groups and maps each group (a value from 0 to 63) to one character of the alphabet. Because 6-bit groups line up with bytes every 24 bits, the work is done in blocks of 3 input bytes and 4 output characters.

A worked example with the text Man:

StepMan
ASCII code7797110
8-bit binary010011010110000101101110

The 24 bits 010011010110000101101110 regroup into four 6-bit values: 010011 (19), 010110 (22), 000101 (5) and 101110 (46). Looking those up in the alphabet gives T, W, F and u, so Man encodes to TWFu. Decoding runs the same steps backwards.

The Base64 alphabet

ValuesCharactersBase64URL difference
0–25A–Zsame
26–51a–zsame
52–610–9same
62+-
63/_
padding=often omitted

All 65 characters are printable ASCII and survive systems that would mangle raw bytes: protocols that only allow 7-bit text, text fields in databases, JSON strings and XML documents.

Padding: what the = signs mean

When the input length is not a multiple of 3, the last block is short. The encoder fills the missing bits with zeros and appends = so that the output length is always a multiple of 4:

  • 3 bytes left: no padding (Man → TWFu).
  • 2 bytes left: one = (Ma → TWE=).
  • 1 byte left: two == (M → TQ==).

Padding carries no data; it only tells the decoder how many bytes the final block holds, and that can also be worked out from the length. That is why many formats drop it. Browsers' atob() accepts input with or without padding (atob("TWE") returns "Ma"), but a string whose length leaves a remainder of 1 when divided by 4 can never be valid and is rejected.

Size overhead

Every 3 bytes become 4 characters, so encoded data is about 33% larger: n bytes produce 4 × ⌈n/3⌉ characters with padding. 100 bytes become 136 characters, and 1,000 bytes become 1,336. MIME email (RFC 2045) also limits encoded lines to 76 characters, each followed by a CRLF line break, which raises the overhead to roughly 37%. PEM files such as certificates wrap at 64 characters. If you send Base64 inside a compressed HTTP response, compression recovers some of the space, but it is still larger than sending the original bytes.

Base64URL (RFC 4648 section 5)

Standard Base64 uses + and /, which have special meanings in URLs (/ separates path segments, and + is read as a space in form-encoded query strings) and / cannot appear in file names. Section 5 of RFC 4648 defines the "URL and filename safe" alphabet, which replaces them with - and _. The same text shows the difference:

Standard:  PDw/Pz8+Pg==   ("<<???>>")
Base64URL: PDw_Pz8-Pg

Where you will meet it:

  • JSON Web Tokens. Each of the three parts of a JWT is Base64URL without padding, defined in RFC 7515, which RFC 7519 builds on. The header {"alg":"HS256","typ":"JWT"} encodes to eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9. Our JWT guide and JWT Decoder cover the rest.
  • URLs and file names: tokens, IDs and signed parameters that must not need percent-encoding.

Mixing the two alphabets is a common bug: a strict standard decoder rejects - and _, and atob("-_-_") throws. Convert first by swapping the characters back and restoring padding if your decoder requires it. In Node.js, Buffer supports both "base64" and "base64url" as encodings.

The btoa() and UTF-8 pitfall in JavaScript

The browser's btoa() ("binary to ASCII") does not take text; it takes a string in which each character stands for one byte, so every character must be in the range U+0000 to U+00FF. Anything outside that range throws an InvalidCharacterError:

btoa("café ✓");   // throws InvalidCharacterError
btoa("café");     // "Y2Fm6Q==" (é encoded as the single Latin-1 byte E9, not UTF-8)

The second case is worse because it silently succeeds: the result is not the UTF-8 encoding that other systems expect (Y2Fmw6k=), and decoding UTF-8 Base64 with plain atob() gives mojibake such as café. Convert to UTF-8 bytes first:

function toBase64(text) {
  const bytes = new TextEncoder().encode(text);
  let binary = "";
  bytes.forEach(b => { binary += String.fromCharCode(b); });
  return btoa(binary);
}

function fromBase64(b64) {
  const bytes = Uint8Array.from(atob(b64), c => c.charCodeAt(0));
  return new TextDecoder().decode(bytes);
}

toBase64("café ✓");          // "Y2Fmw6kg4pyT"
fromBase64("Y2Fmw6kg4pyT");  // "café ✓"

The Base64 Encoder / Decoder works the same way. Encode converts your text to UTF-8 before encoding, so any language or emoji is fine. Decode ignores spaces and line breaks (so wrapped MIME output can be pasted as it is), adds missing padding, accepts the URL-safe characters - and _ even when the option is off, and then requires the decoded bytes to be valid UTF-8. That last step means it decodes Base64 of text, not of binary files: Base64 of an image or a ZIP produces an error instead of garbage. Tick URL-safe (- _, no padding) to produce Base64URL output, with - and _ and the = padding removed.

Data URIs: embedding files in HTML and CSS

A data URI (RFC 2397) puts a file's content directly into a URL, in the form data:[media type][;base64],data:

<img alt="" src="data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI4IiBoZWlnaHQ9IjgiPjxyZWN0IHdpZHRoPSI4IiBoZWlnaHQ9IjgiIGZpbGw9InJlZCIvPjwvc3ZnPg==">

That 106-byte SVG became 144 characters of Base64. The Image to Base64 tool produces this form for you: choose or drop an image and it outputs the complete data URI, with the media type taken from the file, ready to paste into an src attribute or a CSS url(). In the other direction, paste either a full data:image/... URI or raw Base64, pick the MIME type for raw input, and click Show Image to preview it and download it.

Data URIs save a request for small images such as icons, but the encoded data is larger than the file, cannot be cached separately from the page or stylesheet that contains it, and is repeated wherever it is used. For anything beyond a few kilobytes, a normal file is usually the better choice.

Base64 is not encryption

Base64 has no key and no secret. Anyone who sees cGFzc3dvcmQ= can decode it to password in a second. HTTP Basic authentication sends the username and password Base64-encoded, which is why it must only be used over HTTPS, and the payload of a signed JWT is readable by anyone who has the token, so it must not contain secrets. If you need confidentiality, encrypt the data and, if necessary, Base64-encode the ciphertext for transport. If you need to check integrity, use a hash or a signature; see the Hash Generator.

When not to use Base64

  • Large file uploads and downloads. HTTP can carry binary bodies directly (for example multipart/form-data or application/octet-stream); Base64 adds a third to the size and costs memory to encode and decode.
  • Storing binary data in a database that has a binary column type.
  • Hiding values such as API keys or passwords in code or configuration. It hides nothing.
  • Putting text in a URL. For ordinary text, percent-encoding with the URL Encoder keeps it readable; reserve Base64URL for binary values.

Quick reference

  • 3 bytes in, 4 characters out; output length is a multiple of 4 when padded.
  • Size: about 4/3 of the input, or about 37% larger with 76-character MIME line breaks.
  • = means the last block had 2 bytes, == means 1 byte.
  • Base64URL: - instead of +, _ instead of /, padding usually dropped (JWT drops it).
  • In browsers, encode text with TextEncoder before btoa().
  • Never treat Base64 as a security measure.

Frequently asked questions

Why does my Base64 string end with = or ==?

The input length was not a multiple of 3 bytes. One = means the final block held 2 bytes, two mean it held 1 byte. The padding carries no data.

Can I remove the padding?

Only if the decoder accepts unpadded input. Browsers' atob() and the Base64 Encoder / Decoder do, and Base64URL in JWTs is always unpadded, but some strict libraries require the length to be a multiple of 4.

Why does decoding give strange characters such as é?

The bytes were UTF-8 but were decoded as one character per byte, which is what atob() returns. Pass the bytes through TextDecoder, as in the example above.

How much bigger does Base64 make a file?

About a third: 3 bytes become 4 characters, so a 30 KB image becomes about 40 KB of Base64 text, plus a few per cent more if the output is wrapped into lines.

Tools for this guide