Text encoding & Base64
Convert JavaScript strings to UTF-8 bytes, Base64, base64url, and hex with TextEncoder, TextDecoder, atob, btoa, and typed arrays.
- 01Move between strings and bytesUse TextEncoder and TextDecoder without forgetting JavaScript strings are UTF-16 code units.
- 02Read encoded outputExplain UTF-8 byte widths, Base64 padding, base64url changes, and hex pairs from real bytes.
- 03Choose safe browser APIsHandle btoa's binary-string limit, streaming decoders, native typed-array helpers, and practical web formats.
Strings are not bytes
JavaScript lets you write "café" as if it were one simple value, but the browser stores strings as UTF-16 code units. Files, network bodies, cryptographic hashes, images, and typed arrays use bytes. Text encoding is the bridge between those two worlds.
Text encoding turns a JavaScript string into bytes with a named rule, usually UTF-8. Text decoding turns bytes back into a string. Base64 and hex are not character encodings; they are text-friendly ways to print bytes.
A message can be meaningful to a person, but a sorting machine needs a label with exact marks. Encoding writes the label. Decoding scans it back. Base64 and hex are more like tracking numbers: easy to copy through text systems, but still representing the package contents rather than changing them.
- In real life: The message someone wrote
- In JavaScript: The JavaScript string
- In real life: The shipping label format
- In JavaScript: The encoding, such as UTF-8
- In real life: The scanner reading the label
- In JavaScript: The decoder using the same label
- In real life: A tracking number printed for humans
- In JavaScript: Base64 or hex text for bytes
Where the analogy stops: Real shipping labels tolerate smudges and human guesses. Byte encodings are exact: one wrong byte can become �, throw, or decode as the wrong character.
| Representation | Unit | Where it appears |
|---|---|---|
| JavaScript string | UTF-16 code units | Good for text editing; not a byte container. |
| UTF-8 bytes | 1 to 4 bytes per code point | Used by files, fetch bodies, databases, and most web protocols. |
| Base64 | 6-bit text alphabet | Carries bytes through text-only places, usually with = padding. |
| Hex | Two characters per byte | Readable for hashes, packet dumps, colors, and debugging. |
If Unicode terms feel fuzzy, review Unicode & strings. If byte arrays feel new, review the neighboring lesson on ArrayBuffer & typed arrays. Later, the Files & blobs lesson uses the same conversions when reading and writing user files.
UTF-8 writes each code point as 1 to 4 bytes
STEP THROUGHUTF-8 is the web's default byte encoding for text. ASCII characters keep their single byte. Larger code points use leading-byte patterns that announce how many continuation bytes follow. The prefixes are why a decoder can read a stream without spaces between characters.
41c3 a9e2 82 acf0 9f 98 80The next recording implements a tiny UTF-8 encoder for valid code points and proves it agrees with TextEncoder for "Aé€😀". The goal is not to replace the platform API; it is to make the byte patterns inspectable.
Step through a tiny UTF-8 encoder. Watch each character pick the byte pattern that TextEncoder also uses.
script
if (codePoint <= 0x7F) return [codePoint]; if (codePoint <= 0x7FF) { return [0xC0 | (codePoint >> 6), 0x80 | (codePoint & 0x3F)]; } if (codePoint <= 0xFFFF) { return [0xE0 | (codePoint >> 12), 0x80 | ((codePoint >> 6) & 0x3F), 0x80 | (codePoint & 0x3F)]; } return [0xF0 | (codePoint >> 18), 0x80 | ((codePoint >> 12) & 0x3F), 0x80 | ((codePoint >> 6) & 0x3F), 0x80 | (codePoint & 0x3F)];} function encodeUtf8(text) { const bytes = []; for (const character of text) { const codePoint = character.codePointAt(0); bytes.push(...encodeCodePoint(codePoint)); } return new Uint8Array(bytes);} const text = "Aé€😀";console.log([...encodeUtf8(text)].join(","));console.log([...new TextEncoder().encode(text)].join(","));Notice the emoji: JavaScript stores it as two UTF-16 code units, so "😀".length is 2. A for...of loop walks the full code point, and UTF-8 writes four bytes. That is why byte length, code-unit length, and what a person calls “one character” can all differ.
€U+20AC226, 130, 172e2 82 ac€ is U+20AC and encodes as e2 82 ac.
Use TextEncoder and TextDecoder for real work
REAL OUTPUTTextEncoder has one job: turn a string into UTF-8 bytes. It does not accept a label because it always writes UTF-8. encode() creates a new typed array; encodeInto() writes into an existing Uint8Array, which is useful when a parser or stream already owns a buffer.
const encoder = new TextEncoder();const bytes = encoder.encode("éuro");console.log([...bytes].join(",")); const destination = new Uint8Array(5);const result = encoder.encodeInto("éuro", destination);console.log(result.read + " code units");console.log(result.written + " bytes");console.log([...destination].join(","));| API or option | Direction | What to remember |
|---|---|---|
TextEncoder | String to bytes | Always emits UTF-8. Use encodeInto when you already have a destination buffer. |
TextDecoder | Bytes to string | Accepts labels such as utf-8, utf-16le, and windows-1252. |
fatal: true | Reject bad byte sequences | Throws a TypeError instead of inserting U+FFFD. |
stream: true | Decode chunked input | Buffers an unfinished multi-byte character until the next chunk arrives. |
TextDecoder goes from bytes back to strings. It accepts labels such as "utf-8", "utf-16le", and "windows-1252". The label must match the data source, not the page you wish you had. A byte like 0xE9 is a complete é in Windows-1252, but an invalid standalone byte in UTF-8.
const win1252 = new Uint8Array([0x63, 0x61, 0x66, 0xE9]);console.log(new TextDecoder("windows-1252").decode(win1252));console.log(new TextDecoder("utf-8").decode(win1252)); const utf16le = new Uint8Array([0x48, 0x00, 0x69, 0x00]);console.log(new TextDecoder("utf-16le").decode(utf16le)); const withBom = new Uint8Array([0xEF, 0xBB, 0xBF, 0x4F, 0x4B]);console.log(new TextDecoder("utf-8").decode(withBom));console.log(new TextDecoder("utf-8", { ignoreBOM: true }).decode(withBom).charCodeAt(0).toString(16));Bad bytes have two honest outcomes. The default decoder replaces invalid sequences with U+FFFD, shown as �. A decoder created with { fatal: true } throws instead, which is safer for file formats and protocols where corruption should stop the parse.
const bad = new Uint8Array([0xC3, 0x28]);console.log(new TextDecoder("utf-8").decode(bad));console.log(new TextDecoder("utf-8", { fatal: true }).decode(bad));Streams add one more wrinkle: a chunk can end halfway through a multi-byte character. Pass { stream: true } until the final chunk so the decoder can keep the partial byte sequence instead of replacing it too early.
A streaming decoder keeps unfinished multi-byte UTF-8 sequences between chunks. Step through a split é.
script
const first = bytes.slice(0, 4);const second = bytes.slice(4); console.log(new TextDecoder("utf-8").decode(first)); const decoder = new TextDecoder("utf-8", { fatal: true });console.log(decoder.decode(first, { stream: true }));console.log(decoder.decode(second));Base64 is byte text, not Unicode text
EDGE CASESBase64 groups three bytes into four 6-bit indexes. Those indexes pick characters from an alphabet: A-Z, a-z, 0-9, +, and /. Padding = fills the last four-character group when the byte count is not a multiple of three. Base64url swaps + for -, / for _, and often omits padding.
The old browser helpers btoa and atob do not accept arbitrary JavaScript text. They operate on binary strings, where each string character's code is one byte from 0 through 255. That is why this first example works.
console.log(btoa("é"));const binary = atob("6Q==");console.log(binary.charCodeAt(0));The Euro sign is outside that one-byte range, so btoa("€") throws an InvalidCharacterError. The fix is not to guess a code page; encode the text to UTF-8 bytes first.
console.log(btoa("€"));function bytesToBinaryString(bytes) { return String.fromCharCode(...bytes);}function binaryStringToBytes(binary) { return Uint8Array.from(binary, (character) => character.charCodeAt(0));} const bytes = new TextEncoder().encode("€");const base64 = btoa(bytesToBinaryString(bytes));console.log([...bytes].join(","));console.log(base64);console.log(new TextDecoder().decode(binaryStringToBytes(atob(base64))));Anyone can decode Base64 or base64url. Use it when a text-only place must carry bytes, such as JSON, URLs, or a data URL. Do not use it to hide API keys, tokens, or user data.
Modern Uint8Array helpers remove the binary-string detour
FEATURE DETECTThe typed-array Base64 and hex helpers are the direct APIs learners always wanted: bytes in, encoded text out; encoded text in, bytes out. MDN marks toBase64, fromBase64, toHex, and fromHex as standard-track, Baseline 2025 APIs with Chrome 140, Firefox 133, Safari 18.2, and Node 25 support. The CI runtime for this site is Node 22, so examples feature-detect and keep fallbacks.
| API | Shape | Support note |
|---|---|---|
btoa / atob | Binary string only | Old and widely available; every character must be one byte (0 through 255). |
Uint8Array.toBase64() | Bytes to Base64 | Standard-track typed-array helper. MDN marks Baseline 2025: Chrome 140, Firefox 133, Safari 18.2, Node 25. |
Uint8Array.fromBase64() | Base64 to bytes | Supports alphabet: "base64url" and lastChunkHandling; feature-detect because Node 22 lacks it. |
toHex() / fromHex() | Bytes to hex and back | Same Baseline 2025 browser support and Node 25+ in MDN data; use manual hex in Node 22. |
const bytes = new Uint8Array([104, 105, 63]);if (typeof bytes.toBase64 === "function" && typeof Uint8Array.fromBase64 === "function") { console.log(bytes.toBase64()); console.log(bytes.toBase64({ alphabet: "base64url", omitPadding: true })); console.log([...Uint8Array.fromBase64("aGk_", { alphabet: "base64url" })].join(",")); console.log([...Uint8Array.fromBase64("SGk", { lastChunkHandling: "loose" })].join(","));} else { console.log("native Base64 helpers unavailable");} const token = new Uint8Array([222, 173, 190, 239]);if (typeof token.toHex === "function" && typeof Uint8Array.fromHex === "function") { console.log(token.toHex()); console.log([...Uint8Array.fromHex("deadbeef")].join("-"));} else { console.log([...token].map((byte) => byte.toString(16).padStart(2, "0")).join(""));}The Base64 options are worth naming. { alphabet: 'base64url' } switches to URL-safe - and _. { omitPadding: true } drops trailing =. Uint8Array.fromBase64 also accepts lastChunkHandling; "loose" accepts unpadded final chunks, while stricter modes reject or stop before partial input.
Hex is simpler: one byte becomes exactly two base-16 digits. Until native toHex is available everywhere you run, this manual form is easy to audit.
const bytes = new Uint8Array([0, 15, 16, 255]);const hex = [...bytes] .map((byte) => byte.toString(16).padStart(2, "0")) .join("");console.log(hex); const roundTrip = Uint8Array.from(hex.match(/../g), (pair) => parseInt(pair, 16));console.log([...roundTrip].join(","));Recipes you will use on real sites
PLAYGROUNDUse this workbench as a ruler. Type text and compare the UTF-16 code-unit count with code points, UTF-8 bytes, hex, Base64, and base64url. The visible source is intentionally small: encode once, then derive every byte-oriented representation from the same bytes.
const text = "Aé€😀";const bytes = new TextEncoder().encode(text);const hex = [...bytes].map((byte) => byte.toString(16).padStart(2, "0")).join("");const base64 = btoa(String.fromCharCode(...bytes)); console.log([...bytes].join(","));console.log(hex);console.log(base64);5465, 195, 169, 226, 130, 172, 240, 159, 152, 12841c3a9e282acf09f9880QcOp4oKs8J+YgA==QcOp4oKs8J-YgAEvery output starts from the same TextEncoder bytes. Change the text and compare code units, code points, bytes, hex, Base64, and base64url.
JWTs demonstrate why “encoded” does not mean “private.” The middle part is base64url JSON. Decoding it is useful for display and debugging, but a security decision must verify the JWT signature on the server.
function base64UrlToBytes(part) { const base64 = part .replace(/-/g, "+") .replace(/_/g, "/") .padEnd(Math.ceil(part.length / 4) * 4, "="); const binary = atob(base64); return Uint8Array.from(binary, (character) => character.charCodeAt(0));} const payloadPart = "eyJzdWIiOiIxMjMiLCJhZG1pbiI6ZmFsc2V9";const json = new TextDecoder().decode(base64UrlToBytes(payloadPart));console.log(json);console.log(JSON.parse(json).admin);Data URLs are the opposite direction: keep small file bytes inside a string. They are handy for tiny generated files and demos, but large Base64 data URLs bloat HTML and CSS.
const bytes = new TextEncoder().encode("Hello");const base64 = btoa(String.fromCharCode(...bytes));console.log(`data:text/plain;base64,${base64}`);Hashes use bytes too. The Web Crypto API is asynchronous, returns an ArrayBuffer, and leaves the display format to you. Hex is the usual digest display because it is stable, lowercase, and easy to compare.
async function sha256Hex(message) { const bytes = new TextEncoder().encode(message); const digest = await crypto.subtle.digest("SHA-256", bytes); return [...new Uint8Array(digest)] .map((byte) => byte.toString(16).padStart(2, "0")) .join("");} sha256Hex("hello").then(console.log);for (const length of [1, 2, 3, 4, 12]) { const base64Length = Math.ceil(length / 3) * 4; console.log(length + " bytes -> " + base64Length + " Base64 chars");}new TextEncoder().encode(message)new TextDecoder('utf-8', { fatal: true })JWT payload part: eyJzdWIiOiIxMjMifQdata:image/png;base64,...2cf24dba5fb0a30e...byte.toString(16).padStart(2, '0')
Sort each real-world snippet by the conversion or byte display it needs.
Common misconceptions
- “A JavaScript string is already UTF-8.” It is UTF-16 code units. UTF-8 appears after you encode it.
- “
btoaencodes Unicode text.” It encodes binary strings. UseTextEncoderfor arbitrary text. - “Base64 makes data safe.” It only makes bytes text-safe. It is not encryption, signing, escaping, or validation.
- “A decoder always throws on bad bytes.” The default inserts
�. Add{ fatal: true }when invalid input should stop. - “Hex and Base64 are interchangeable.” Both display bytes, but hex is two characters per byte while Base64 is denser and chunked.
| Confusing idea | Reality | Safer habit |
|---|---|---|
btoa(text) | Only safe for binary strings | Encode text with TextEncoder first, then Base64 the bytes. |
| Base64 | A transport encoding | It is reversible without a secret key, so it is not encryption. |
| Hex | A debugging-friendly encoding | It doubles size, but every two characters map to exactly one byte. |
TextDecoder default | Replacement mode | Bad bytes become �; add fatal: true when corrupt input should fail. |
Practice exercises
5 EXERCISESRead the snippet and type the exact decimal bytes printed by the console.
const bytes = new TextEncoder().encode("é");
console.log([...bytes].join(","));TextEncoder writes é as bytes 0xC3 and 0xA9, which print as decimal 195,169.
Predict the hex string. This checks why padStart(2, "0") matters.
const bytes = new Uint8Array([15, 16, 255]);
const hex = [...bytes]
.map((byte) => byte.toString(16).padStart(2, "0"))
.join("");
console.log(hex);The padded pairs are 0f, 10, and ff, so the joined hex string is 0f10ff.
The code throws. Type the browser API that should run before btoa for arbitrary Unicode text.
console.log(btoa("€"));const bytes = new TextEncoder().encode("€");
const binary = String.fromCharCode(...bytes);
console.log(btoa(binary));TextEncoder turns € into UTF-8 bytes 226,130,172; those bytes can be converted to a binary string and passed to btoa.
Type the boolean printed by the final line.
function base64UrlToBytes(part) {
const base64 = part.replace(/-/g, "+").replace(/_/g, "/").padEnd(Math.ceil(part.length / 4) * 4, "=");
return Uint8Array.from(atob(base64), (character) => character.charCodeAt(0));
}
const json = new TextDecoder().decode(base64UrlToBytes("eyJzdWIiOiIxMjMiLCJhZG1pbiI6ZmFsc2V9"));
console.log(JSON.parse(json).admin);The payload decodes to {"sub":"123","admin":false}, so JSON.parse(json).admin prints false.
A payload has 5 bytes before encoding. How many Base64 characters will represent it?
const length = 5;
console.log(Math.ceil(length / 3) * 4);Math.ceil(5 / 3) is 2 chunks, and each chunk is 4 Base64 characters, so the output length is 8.
Check your understanding
7 QUESTIONSQuestion 1 of 7Which statement best describes JavaScript strings versus files or network data?
Choose an answer to see the explanation.
Question 2 of 7What bytes does UTF-8 use for
€?Read the code, then predictconst bytes = new TextEncoder().encode("€"); console.log([...bytes].join(","));Choose an answer to see the explanation.
Question 3 of 7What does replacement-mode decoding print for invalid UTF-8?
Read the code, then predictconst bad = new Uint8Array([0xC3, 0x28]); console.log(new TextDecoder("utf-8").decode(bad));Choose an answer to see the explanation.
Question 4 of 7What does
btoa('é')print in a browser?Read the code, then predictconsole.log(btoa("é"));Choose an answer to see the explanation.
Question 5 of 7Which Base64 form is URL-safe and omits padding?
Read the code, then predictconst base64 = "aGk/"; console.log(base64.replace(/\+/g, "-").replace(/\//g, "_").replace(/=+$/g, ""));Choose an answer to see the explanation.
Question 6 of 7What hex string does the manual encoder produce?
Read the code, then predictconst bytes = new Uint8Array([0, 15, 255]); console.log([...bytes].map((byte) => byte.toString(16).padStart(2, "0")).join(""));Choose an answer to see the explanation.
Question 7 of 7Why is Base64 about one third larger than the original bytes?
Read the code, then predictfor (const length of [3, 6]) { console.log(Math.ceil(length / 3) * 4); }Choose an answer to see the explanation.
Key takeaways
- JavaScript strings are UTF-16 code units; most storage and transport APIs work with bytes.
TextEncoderalways writes UTF-8;TextDecoderreads bytes with labels and options.- UTF-8 uses one to four bytes per code point; byte length is not string length.
btoaandatobuse binary strings, so encode Unicode text to bytes first.- Base64 and hex represent bytes as text. They are reversible encodings, not encryption.
- Native
Uint8ArrayBase64 and hex helpers are useful, but feature-detect for Node 22 and older browsers.
Remember the one-liner.
Encode text before treating it as bytes, and label byte display formats like Base64 and hex for what they are: transport and debugging tools, not secrecy.
Up next: memory management and garbage collection, where the lifetime of objects, buffers, and references becomes the main story.