Utilo

Base64 Explained: How It Works, When to Use It and When Not To

A clear explanation of Base64 encoding: the algorithm, padding, URL-safe Base64, data URLs, size overhead, Unicode pitfalls and security misconceptions.

· 5 min read

Base64 shows up in email attachments, JSON Web Tokens, data URLs, Kubernetes secrets, HTTP Basic authentication and countless APIs that need to move binary data through text-only channels. It is simple once you see how it works, yet it is often misused — most notably as if it were encryption. This article explains the algorithm step by step, the common variants, and practical guidance on when Base64 is the right tool.

The problem Base64 solves

Many systems were designed to carry text, not arbitrary bytes. Email (SMTP) historically supported only 7-bit ASCII. JSON strings cannot contain raw binary. URLs and HTTP headers have restricted character sets. If you put an image's raw bytes into any of these, some bytes will be interpreted as control characters, line endings or delimiters, and the data gets corrupted.

Base64 maps arbitrary bytes onto 64 safe, printable characters: A–Z, a–z, 0–9, + and /. Any system that can carry plain text can then carry the data intact.

How the encoding works

Base64 processes input in groups of 3 bytes (24 bits) and outputs 4 characters of 6 bits each. Since 2^6 = 64, each 6-bit value selects one character from the alphabet.

Take the word Man:

  1. ASCII bytes: M = 77, a = 97, n = 110.
  2. In binary: 01001101 01100001 01101110.
  3. Split into four 6-bit groups: 010011 010110 000101 101110.
  4. As numbers: 19, 22, 5, 46.
  5. Look up in the alphabet: T, W, F, u.

So Man becomes TWFu. Decoding reverses the process.

Padding with =

When the input length is not a multiple of 3, the last group is incomplete. Base64 pads it:

  • 1 leftover byte produces 2 characters plus ==.
  • 2 leftover bytes produce 3 characters plus =.

Ma encodes to TWE= and M to TQ==. Padding makes the output length a multiple of four, which some decoders require. Others, especially URL-safe implementations, omit it because the length already implies how many bytes are missing.

Size overhead

Every 3 bytes become 4 characters, so Base64 output is about 33% larger than the input, plus padding and sometimes line breaks (MIME email wraps lines at 76 characters). A 1 MB image becomes roughly 1.37 MB of text. That overhead matters when you embed large files in JSON or HTML.

URL-safe Base64

The standard alphabet includes + and /, which have special meanings in URLs, and = which is used in query strings. Base64URL, defined in RFC 4648, replaces + with - and / with _, and usually drops padding. JWTs, many API tokens and web push keys use this variant — see how JWT works for an example.

A standard decoder fails on - and _, and vice versa, which is a common source of "invalid character" errors. A tolerant Base64 decoder accepts both alphabets and restores missing padding automatically.

Text, bytes and Unicode

Base64 encodes bytes, not characters. To encode text, you must first convert it to bytes with a character encoding, almost always UTF-8. This is where browsers trip people up: JavaScript's built-in btoa() accepts only characters in the Latin-1 range, so btoa("héllo") works by accident and btoa("你好") throws an error.

The correct approach in modern JavaScript:

const bytes = new TextEncoder().encode("你好 👋");
const b64 = btoa(String.fromCharCode(...bytes));
const back = new TextDecoder().decode(Uint8Array.from(atob(b64), c => c.charCodeAt(0)));

Other languages make this explicit: Python's base64.b64encode takes bytes, so you write base64.b64encode("你好".encode("utf-8")).

Data URLs

A data URL embeds a file directly in HTML or CSS:

<img src="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAA..." alt="">

This saves an HTTP request, which was valuable in the HTTP/1.1 era. Today, with HTTP/2 multiplexing, the trade-offs usually favour separate files:

  • Data URLs cannot be cached independently; they are re-downloaded with every page or stylesheet that contains them.
  • They inflate HTML and CSS, which blocks rendering until downloaded and parsed.
  • The 33% overhead applies, although gzip recovers some of it.

Use data URLs for tiny assets — small icons under 1–2 KB, placeholder images, or single-file HTML exports. For anything larger, serve a normal file and compress it instead.

Base64 is not encryption

This is the most important point. Base64 has no key. Anyone can decode it instantly, and many tools do it automatically. Yet it is common to find:

  • API keys "hidden" in Base64 in front-end code.
  • Passwords stored Base64-encoded in databases or config files.
  • Kubernetes Secrets assumed to be protected because their values are Base64.

Kubernetes encodes secret values in Base64 purely so binary data fits in YAML; protection comes from access control and encryption at rest, which must be configured separately. If data needs to be confidential, encrypt it with a real algorithm such as AES-GCM, and if you need to verify passwords, use a password hashing function — not Base64 and not a plain hash like MD5 (why not).

Common use cases done right

  • HTTP Basic authentication sends Authorization: Basic base64(user:password). The encoding only avoids issues with special characters; the credentials are protected solely by HTTPS.
  • Email attachments use MIME Base64 with line breaks every 76 characters.
  • Embedding binary in JSON, such as a small signature image or a cryptographic key, is a legitimate use. For large files, upload separately and reference them by URL.
  • Storing binary in text-only configuration, such as TLS certificates in environment variables.

Troubleshooting decoding errors

  • Invalid character: you are probably decoding Base64URL with a standard decoder, or the string contains whitespace or line breaks.
  • Incorrect padding: add = until the length is a multiple of four.
  • Garbled text after decoding: the original was encoded from a different character set, or the data is binary rather than text.
  • Output looks like Base64 again: the data was encoded twice; decode once more.
  • Base32 uses 32 characters (A–Z and 2–7), is case-insensitive and is used in TOTP secrets for authenticator apps. Overhead is 60%.
  • Hexadecimal (Base16) uses two characters per byte — 100% overhead — but is easy to read and common for hashes.
  • Base58 removes look-alike characters and is used in Bitcoin addresses.
  • Percent-encoding escapes only unsafe characters in URLs.

Summary

Base64 converts bytes to a 64-character text alphabet, 3 bytes at a time, at a cost of about 33% extra size. Use it to carry binary data through text channels; use the URL-safe variant in URLs and tokens; always convert text to UTF-8 bytes first. Never use it to protect secrets — it is an encoding, not a lock.

Related guides