Base64 is encoding, not encryption. Anyone with the string can decode it back to the original content.
Hello -> SGVsbG8=Base64 converts binary bytes into text made from common ASCII characters, which makes it convenient for JSON, HTML, CSS, configs, and API fields.
How Encoding Works: 3 Bytes Become 4 Characters
Each byte contains 8 bits. Base64 takes 3 bytes, or 24 bits, and splits them from left to right into four 6-bit groups. Each group represents a number from 0 to 63, which selects one character from a 64-character alphabet. That alphabet gives Base64 its name.
| Index (starting at 0) | Standard Base64 character |
|---|---|
| 0–25 | A–Z |
| 26–51 | a–z |
| 52–61 | 0–9 |
| 62 | + |
| 63 | / |
For example, each character in Man occupies one byte in both ASCII and UTF-8. Write those bytes in binary, regroup them into 6-bit values, and look up each value to get TWFu. This changes the representation of the data without using a key or compressing it.
Man → TWFu
Characters: M a n
Decimal bytes: 77 97 110
8-bit groups: 01001101 01100001 01101110
6-bit groups: 010011 010110 000101 101110
Indexes: 19 22 5 46
Alphabet: T W F uWhy Pad with = When Fewer Than 3 Bytes Remain?
If the final group contains fewer than 3 bytes, append zero bits on the right to complete the last 6-bit value, then add = characters to make four output characters. The = sign is a padding marker outside the 64-character alphabet; it does not represent an equals sign or a zero byte in the original data.
| Input | 6-bit groups (last group zero-padded) | Data indexes | Padded output |
|---|---|---|---|
| M (1 byte) | 010011 010000 | 19, 16 | TQ== |
| Ma (2 bytes) | 010011 010110 000100 | 19, 22, 4 | TWE= |
| Man (3 bytes) | 010011 010110 000101 101110 | 19, 22, 5, 46 | TWFu |
Decoding reverses the process: map characters back to 6-bit values, join the bits, and recover 8-bit bytes, discarding the trailing bits added during encoding as indicated by the padding. TQ== therefore restores 01001101, or M. Some protocols allow omitted padding, but both sides must agree on that convention.
Why Does Base64 Get Larger?
For n input bytes, padded Base64 contains 4 × ceil(n / 3) ASCII characters, where ceil rounds up, excluding line breaks and any Data URL prefix. For large inputs the overhead approaches one third; short inputs can have a higher ratio, such as 1 byte becoming 4 characters.
Data URL vs Raw Base64
| Form | Example | Use case |
|---|---|---|
| Raw Base64 | iVBORw0KGgo... | Only the encoded content |
| Data URL | data:image/png;base64,iVBOR... | Includes MIME type and can be used as img src |
| URL Safe Base64 | Uses - and _ | Common in JWT and URL contexts |
Common Pitfalls
- Treating Base64 as encryption and exposing sensitive data.
- Missing trailing padding characters when copying.
- Sending a Data URL when the backend expects raw Base64.
- Embedding very large images as Base64 and slowing down pages or APIs.
- Mixing standard Base64 and URL Safe Base64.
Why text still needs a character encoding
Base64 encodes bytes, not abstract characters. English Hello is five bytes in UTF-8, while Chinese text, emoji, and other characters first become different UTF-8 byte sequences and are then encoded. If the producer and consumer use different character sets, the Base64 can be perfectly valid while the restored text is still corrupted.
text -> UTF-8 bytes -> Base64
Base64 -> bytes -> decode with the same character setWhen it fits and when it does not
| Situation | Recommendation | Reason |
|---|---|---|
| A very small icon in CSS | Consider it | It removes one request but enlarges the text asset |
| A small binary field in JSON | Confirm the API contract | MIME type, length, and size limits need an explicit convention |
| A large image or video | Do not inline it | Expansion, memory use, and parsing cost all become significant |
| Passwords, tokens, or personal data | Never treat it as protection | No key is needed; anyone with the string can restore it |
When decoding fails, first check for an included Data URL prefix, spaces, or line breaks, then confirm whether the alphabet is standard or URL-safe. Some protocols omit trailing padding deliberately, but the receiver must support that convention. Do not guess at damaged content by adding equals signs blindly.