What Is Base64? How Base64 Encoding Works and When to Use It
Base64 is a way to write any binary data, such as an image, a PDF, or a cryptographic key, using only 64 ordinary text characters: the letters A–Z and a–z, the digits 0–9, and the symbols + and /. It splits the data into 6-bit chunks and maps each chunk to one of those characters. That lets binary content pass safely through systems built to carry text, such as email, JSON, and HTTP headers.
The trade-off is size. Base64 output is about 33% larger than the input, because every 3 bytes become 4 characters. It also isn't encryption. Anyone can decode Base64 instantly, without a key.
This guide encodes a real example bit by bit and explains the = padding and the URL-safe variant. It also covers where Base64 shows up in everyday tech and the mistakes that most often break it.
How Base64 encoding works
A byte holds 8 bits. A Base64 character carries 6 bits, because 2⁶ = 64. Three bytes give 24 bits, and 24 bits split evenly into four 6-bit groups. So Base64 always works in blocks of 3 input bytes and 4 output characters.
The 64-character alphabet
Each 6-bit group is a number from 0 to 63. The standard alphabet, defined in RFC 4648, maps those numbers like this:
| 6-bit value | Character | URL-safe version |
|---|---|---|
| 0–25 | A–Z | same |
| 26–51 | a–z | same |
| 52–61 | 0–9 | same |
| 62 | + |
- |
| 63 | / |
_ |
| (padding) | = |
often omitted |
To find a character, subtract the start of its range. For example, value 46 falls in the lowercase range, and 46 − 26 = 20, which is the 21st lowercase letter, u.
Worked example: encoding "Man" to "TWFu"
The text "Man" is three bytes in ASCII, a perfect single block.
- Convert each character to its byte value. M = 77, a = 97, n = 110.
- Write each byte as 8 bits. 77 =
01001101, 97 =01100001, 110 =01101110. - Join the bits into one 24-bit string.
010011010110000101101110 - Split it into four 6-bit groups.
010011010110000101101110 - Convert each group to a number.
010011= 16 + 2 + 1 = 19.010110= 16 + 4 + 2 = 22.000101= 4 + 1 = 5.101110= 32 + 8 + 4 + 2 = 46. - Look up each number in the alphabet. 19 =
T, 22 =W, 5 =F, 46 =u.
Result: "Man" encodes to TWFu. Decoding runs the same steps in reverse. Each character becomes its 6-bit value, and the bits are regrouped into bytes of 8.
Base64 never looks at what the bytes mean. Text, a JPEG, or a ZIP file all go through exactly the same steps.
What the "=" padding at the end means
Input doesn't always come in multiples of 3 bytes. When the last block has only 1 or 2 bytes, the encoder fills the missing bits with zeros. It then adds = signs so the output still comes in groups of 4 characters.
- "Ma" (2 bytes, 16 bits):
01001101 01100001becomes0100110101100001+00. That gives 19, 22, 4, which isTWE, plus one=. Result:TWE=. - "M" (1 byte, 8 bits):
01001101becomes01001101+0000. That gives 19, 16, which isTQ, plus two=. Result:TQ==.
So the padding tells you how many bytes the last block held. No = means a full 3 bytes, one = means 2 bytes, and two = means 1 byte. You will never see three = signs in valid Base64, because a single leftover byte still produces two real characters.
Padding is technically redundant, since the decoder can work out the leftover bytes from the string length. That's why some formats drop it. A padded string's length is always a multiple of 4, though, and some strict decoders require that.
Why Base64 output is about 33% larger
Every 3 bytes become 4 characters, so the output is 4/3 the size of the input, about 133%. The exact length of standard padded Base64 for n input bytes is:
output length = 4 × ⌈n ÷ 3⌉ (⌈ ⌉ means round up to the next whole number)
The padding depends on the remainder when n is divided by 3: remainder 0 means no padding, remainder 2 means =, and remainder 1 means ==.
| Input bytes | Remainder (n ÷ 3) | Calculation | Output characters | Padding | Growth |
|---|---|---|---|---|---|
| 1 | 1 | 4 × ⌈1/3⌉ = 4 × 1 | 4 | == |
+300% |
| 2 | 2 | 4 × ⌈2/3⌉ = 4 × 1 | 4 | = |
+100% |
| 3 | 0 | 4 × ⌈3/3⌉ = 4 × 1 | 4 | none | +33% |
| 10 | 1 | 4 × ⌈10/3⌉ = 4 × 4 | 16 | == |
+60% |
| 100 | 1 | 4 × ⌈100/3⌉ = 4 × 34 | 136 | == |
+36% |
| 101 | 2 | 4 × ⌈101/3⌉ = 4 × 34 | 136 | = |
+34.7% |
| 1,000 | 1 | 4 × ⌈1,000/3⌉ = 4 × 334 | 1,336 | == |
+33.6% |
| 1,000,000 | 1 | 4 × ⌈1,000,000/3⌉ = 4 × 333,334 | 1,333,336 | == |
+33.3% |
For tiny inputs the overhead looks large because of padding. Notice that 100 and 101 bytes both produce 136 characters: the 101st byte fills a slot that padding would otherwise take. For anything over a few hundred bytes the overhead settles at about 33%. Line breaks add a little more. Email wraps Base64 into lines of at most 76 characters, which pushes the overhead a few points higher.
To go the other way, estimate the decoded size as characters × 3 ÷ 4, minus one byte for each =. For example, TWE= is 4 × 3 ÷ 4 − 1 = 2 bytes.
Base64 vs. Base64URL: the URL-safe variant
Two characters in standard Base64 cause trouble on the web. / separates path segments in URLs and file paths. + is read as a space in form-encoded query strings. A third character, =, separates keys from values in query strings.
RFC 4648 defines a second alphabet, Base64URL, for these situations:
+becomes-/becomes_=padding is usually dropped when the length is known from context
Everything else is identical, so the same bytes give nearly the same string in both. For example, the bytes FB FF encode to +/8= in standard Base64 and -_8 in Base64URL without padding.
You'll find Base64URL in JSON Web Tokens (JWTs). Signed JWTs, the common kind, are three Base64URL segments separated by dots. Encrypted JWTs use a different, five-part format. Base64URL is also common in URL parameters, filenames, and API tokens.
Don't assume the two variants are interchangeable. A strict standard decoder will reject - and _. Base64URL decoders vary by library: some reject + and /, while others accept them. Converting to the alphabet the decoder expects is the safe approach.
Base64 vs. other binary-to-text encodings
Base64 isn't the only binary-to-text encoding. It's popular because it sits in a practical middle ground:
| Encoding | Characters used | Bytes → characters | Size overhead |
|---|---|---|---|
| Hexadecimal (Base16) | 0–9, A–F | 1 → 2 | 100% |
| Base32 | A–Z, 2–7 | 5 → 8 | 60% |
| Base64 | A–Z, a–z, 0–9, +, / | 3 → 4 | ~33% |
| Ascii85 | 85 printable ASCII characters | 4 → 5 | 25% |
Hex is easiest for humans to read. Base32 avoids lowercase and look-alike characters, so it suits codes people type by hand. Ascii85 is more compact but uses characters such as quotes and backslashes that often need escaping. Base64's alphabet avoids most of those problems, so it's the usual choice.
Where Base64 is used
Data URLs in HTML and CSS
A data URL embeds a file directly in a web page instead of linking to it. The format is data:[media type][;base64],[data]. For example, data:text/plain;base64,SGVsbG8= is the text "Hello". Small icons are often inlined this way to save a separate network request.
The 33% growth matters here. A 900-byte SVG icon becomes 4 × ⌈900 ÷ 3⌉ = 4 × 300 = 1,200 characters, 300 bytes more. A 300,000-byte photo becomes 400,000 characters. Inlining a large image also stops the browser from caching it separately, so anything beyond small icons is usually better as a regular image file. SVG is already text, so it can go into a data URL without Base64 if its special characters are percent-encoded.
Email attachments (MIME)
Email was designed for plain text, and older mail servers weren't guaranteed to pass arbitrary bytes intact. The MIME standard solves this by Base64-encoding attachments. It marks them with the header Content-Transfer-Encoding: base64 and wraps the text into lines of at most 76 characters. That's why a 3 MB attachment makes the message noticeably larger than 3 MB, and why mailbox size limits fill faster than file sizes suggest.
JSON and API payloads
JSON has no binary data type, so its strings must hold text. When an API needs to send an image, a file, or a key inside JSON, it typically Base64-encodes the bytes into a string field. It's simple and works everywhere, but you pay the 33% overhead plus the cost of encoding and decoding. For large files, many APIs offer a separate upload method instead.
HTTP Basic authentication
HTTP Basic auth, defined in RFC 7617, sends credentials in a header. The client joins the username and password with a colon and Base64-encodes the result. RFC 7617's own example uses user "Aladdin" and password "open sesame":
Authorization: Basic QWxhZGRpbjpvcGVuIHNlc2FtZQ==
Base64 is used here so that any characters in the password fit safely in a header line. It isn't there for security, which leads to the most important point in this article.
Base64 is not encryption
Base64 has no key and no secret. Anyone who sees QWxhZGRpbjpvcGVuIHNlc2FtZQ== can decode it to Aladdin:open sesame in a second, in any browser console or terminal. Encoded data only looks scrambled.
What this means in practice:
- Never use Base64 to "hide" passwords, API keys, or personal data in code, config files, URLs, or databases.
- HTTP Basic auth is only safe over HTTPS, because HTTPS encrypts the whole request, header included.
- Signed JWT contents are readable by anyone. The signature stops tampering, but the data in the token isn't confidential unless the token is encrypted.
- Base64 also isn't hashing or compression. It's fully reversible, and it makes data bigger, not smaller.
If you need confidentiality, use real encryption such as TLS for data in transit or an established encryption library for stored data. Base64 may still appear at the end, to turn the encrypted bytes into text.
Common Base64 mistakes and how to fix them
Mixing the two alphabets. Decoding a JWT segment with a strict standard decoder fails on - or _. Fix it by replacing - with + and _ with /. Then add = until the length is a multiple of 4.
Missing padding. Some decoders reject a string like TWE but accept TWE=. Add = signs to reach a multiple of 4. Never add more than two. If the length leaves a remainder of 1 when divided by 4, the string itself is damaged.
+ turning into a space. If standard Base64 travels in a URL query string, a + may arrive as a space. Use Base64URL instead, or percent-encode the value before putting it in the URL.
Different results for the same text. Base64 encodes bytes, not characters, so the text encoding matters. "é" is 2 bytes in UTF-8 (C3 A9), which encodes to w6k=. In Latin-1 it's 1 byte (E9), which encodes to 6Q==. Agree on UTF-8 at both ends. In browser JavaScript, btoa() accepts only characters in the Latin-1 range, so convert other text to UTF-8 bytes first.
Line breaks and whitespace. Some tools wrap output every 76 characters (MIME style), and some decoders reject the line breaks. Many command-line base64 tools have an option to turn wrapping off. Check your tool's documentation, because the flag varies by version and operating system.
Double encoding. If the decoded result looks like Base64 again, something encoded the data twice. Decode one more time, then remove the extra encoding step at its source.
Base64 quick reference
- Alphabet: A–Z, a–z, 0–9,
+,/; padding=. - URL-safe alphabet:
-and_replace+and/; padding usually dropped. - Block size: 3 bytes in, 4 characters out.
- Encoded length: 4 × ⌈bytes ÷ 3⌉ characters with padding.
- Padding from input size: bytes ÷ 3 leaves remainder 0 → none, 2 →
=, 1 →==. Three=is never valid. - Decoded size: characters × 3 ÷ 4, minus 1 per
=. - Overhead: about 33%, a bit more with MIME line breaks.
- Good uses: small inline images, email attachments, binary fields in JSON, header values.
- Bad uses: hiding secrets, shrinking data, embedding large files in web pages.
- Before decoding a failing string: check which alphabet it uses, fix the padding, strip whitespace, and confirm the text encoding.
Frequently Asked Questions
How can I tell if a string is Base64?
Look for a string that uses only A–Z, a–z, 0–9, + and / (or - and _ for the URL-safe form), possibly ending in one or two = signs, with a length that's a multiple of 4 when padded. Those signs are suggestive, not proof, because ordinary words like 'Test' also fit the pattern. The only real check is to decode it and see whether the result makes sense.
Does Base64 compress data?
No. Base64 always makes data about a third larger. If you need both compression and text-safe output, compress the data first (for example with gzip) and then Base64-encode the compressed bytes. Doing it the other way around works poorly.
Is Base64 the same as hashing?
No. A hash such as SHA-256 is a one-way fingerprint: you can't get the original data back, and the output is a fixed length. Base64 is fully reversible, and its output grows with the input. Hash values are sometimes written in Base64 or hex, but the encoding is only the display format.
Why is it called Base64?
The name comes from the 64 symbols used to represent values, just as base 10 uses ten digits and base 16 (hexadecimal) uses sixteen symbols. Sixty-four is 2 to the 6th power, so each character stands for exactly 6 bits. That's what makes 3 bytes fit evenly into 4 characters.
Can Base64 strings contain spaces or line breaks?
The encoded data itself never includes spaces, but some formats insert line breaks for transport. MIME email, for example, wraps lines at 76 characters. Many decoders ignore this whitespace, but strict ones reject it, so remove it before decoding if you get an error.