cf.completefrontendCode editorOpen lab
THE JAVASCRIPT FIELD GUIDE

Text encoding & Base64

Convert JavaScript strings to UTF-8 bytes, Base64, base64url, and hex with TextEncoder, TextDecoder, atob, btoa, and typed arrays.

By the end, you can
  • 01
    Move between strings and bytesUse TextEncoder and TextDecoder without forgetting JavaScript strings are UTF-16 code units.
  • 02
    Read encoded outputExplain UTF-8 byte widths, Base64 padding, base64url changes, and hex pairs from real bytes.
  • 03
    Choose safe browser APIsHandle btoa's binary-string limit, streaming decoders, native typed-array helpers, and practical web formats.

Strings are not bytes

JavaScript lets you write "café" as if it were one simple value, but the browser stores strings as UTF-16 code units. Files, network bodies, cryptographic hashes, images, and typed arrays use bytes. Text encoding is the bridge between those two worlds.

Definition

Text encoding turns a JavaScript string into bytes with a named rule, usually UTF-8. Text decoding turns bytes back into a string. Base64 and hex are not character encodings; they are text-friendly ways to print bytes.

Real-life analogyA message, a shipping label, and a scanner

A message can be meaningful to a person, but a sorting machine needs a label with exact marks. Encoding writes the label. Decoding scans it back. Base64 and hex are more like tracking numbers: easy to copy through text systems, but still representing the package contents rather than changing them.

In real life: The message someone wrote
In JavaScript: The JavaScript string
In real life: The shipping label format
In JavaScript: The encoding, such as UTF-8
In real life: The scanner reading the label
In JavaScript: The decoder using the same label
In real life: A tracking number printed for humans
In JavaScript: Base64 or hex text for bytes

Where the analogy stops: Real shipping labels tolerate smudges and human guesses. Byte encodings are exact: one wrong byte can become �, throw, or decode as the wrong character.

Four representations you will switch between
RepresentationUnitWhere it appears
JavaScript stringUTF-16 code unitsGood for text editing; not a byte container.
UTF-8 bytes1 to 4 bytes per code pointUsed by files, fetch bodies, databases, and most web protocols.
Base646-bit text alphabetCarries bytes through text-only places, usually with = padding.
HexTwo characters per byteReadable for hashes, packet dumps, colors, and debugging.

If Unicode terms feel fuzzy, review Unicode & strings. If byte arrays feel new, review the neighboring lesson on ArrayBuffer & typed arrays. Later, the Files & blobs lesson uses the same conversions when reading and writing user files.

UTF-8 writes each code point as 1 to 4 bytes

STEP THROUGH

UTF-8 is the web's default byte encoding for text. ASCII characters keep their single byte. Larger code points use leading-byte patterns that announce how many continuation bytes follow. The prefixes are why a decoder can read a stream without spaces between characters.

AU+00411 byte41
éU+00E92 bytesc3 a9
€U+20AC3 bytese2 82 ac
😀U+1F6004 bytesf0 9f 98 80

The next recording implements a tiny UTF-8 encoder for valid code points and proves it agrees with TextEncoder for "Aé€😀". The goal is not to replace the platform API; it is to make the byte patterns inspectable.

UTF-8 pattern lab
Step 0 of 9Ready
Your turn: follow the blue line

Step through a tiny UTF-8 encoder. Watch each character pick the byte pattern that TextEncoder also uses.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
  if (codePoint <= 0x7F) return [codePoint];  if (codePoint <= 0x7FF) {    return [0xC0 | (codePoint >> 6),            0x80 | (codePoint & 0x3F)];  }  if (codePoint <= 0xFFFF) {    return [0xE0 | (codePoint >> 12),            0x80 | ((codePoint >> 6) & 0x3F),            0x80 | (codePoint & 0x3F)];  }  return [0xF0 | (codePoint >> 18),          0x80 | ((codePoint >> 12) & 0x3F),          0x80 | ((codePoint >> 6) & 0x3F),          0x80 | (codePoint & 0x3F)];} function encodeUtf8(text) {  const bytes = [];  for (const character of text) {    const codePoint = character.codePointAt(0);    bytes.push(...encodeCodePoint(codePoint));  }  return new Uint8Array(bytes);} const text = "Aé€😀";console.log([...encodeUtf8(text)].join(","));console.log([...new TextEncoder().encode(text)].join(","));
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

Notice the emoji: JavaScript stores it as two UTF-16 code units, so "😀".length is 2. A for...of loop walks the full code point, and UTF-8 writes four bytes. That is why byte length, code-unit length, and what a person calls “one character” can all differ.

Inspect one character
Character€
Code pointU+20AC
Bytes226, 130, 172
Hexe2 82 ac
Try it yourself

€ is U+20AC and encodes as e2 82 ac.

This small inspector uses the same manual encoder as the step-through.

Use TextEncoder and TextDecoder for real work

REAL OUTPUT

TextEncoder has one job: turn a string into UTF-8 bytes. It does not accept a label because it always writes UTF-8. encode() creates a new typed array; encodeInto() writes into an existing Uint8Array, which is useful when a parser or stream already owns a buffer.

TextEncoder and encodeIntoPop out in the code editor (opens in a new tab)JavaScript
const encoder = new TextEncoder();const bytes = encoder.encode("éuro");console.log([...bytes].join(",")); const destination = new Uint8Array(5);const result = encoder.encodeInto("éuro", destination);console.log(result.read + " code units");console.log(result.written + " bytes");console.log([...destination].join(","));
Text encoding API choices
API or optionDirectionWhat to remember
TextEncoderString to bytesAlways emits UTF-8. Use encodeInto when you already have a destination buffer.
TextDecoderBytes to stringAccepts labels such as utf-8, utf-16le, and windows-1252.
fatal: trueReject bad byte sequencesThrows a TypeError instead of inserting U+FFFD.
stream: trueDecode chunked inputBuffers an unfinished multi-byte character until the next chunk arrives.

TextDecoder goes from bytes back to strings. It accepts labels such as "utf-8", "utf-16le", and "windows-1252". The label must match the data source, not the page you wish you had. A byte like 0xE9 is a complete é in Windows-1252, but an invalid standalone byte in UTF-8.

Labels and the byte order markPop out in the code editor (opens in a new tab)JavaScript
const win1252 = new Uint8Array([0x63, 0x61, 0x66, 0xE9]);console.log(new TextDecoder("windows-1252").decode(win1252));console.log(new TextDecoder("utf-8").decode(win1252)); const utf16le = new Uint8Array([0x48, 0x00, 0x69, 0x00]);console.log(new TextDecoder("utf-16le").decode(utf16le)); const withBom = new Uint8Array([0xEF, 0xBB, 0xBF, 0x4F, 0x4B]);console.log(new TextDecoder("utf-8").decode(withBom));console.log(new TextDecoder("utf-8", { ignoreBOM: true }).decode(withBom).charCodeAt(0).toString(16));

Bad bytes have two honest outcomes. The default decoder replaces invalid sequences with U+FFFD, shown as �. A decoder created with { fatal: true } throws instead, which is safer for file formats and protocols where corruption should stop the parse.

Replacement versus fatal decodingPop out in the code editor (opens in a new tab)JavaScript
const bad = new Uint8Array([0xC3, 0x28]);console.log(new TextDecoder("utf-8").decode(bad));console.log(new TextDecoder("utf-8", { fatal: true }).decode(bad));

Streams add one more wrinkle: a chunk can end halfway through a multi-byte character. Pass { stream: true } until the final chunk so the decoder can keep the partial byte sequence instead of replacing it too early.

Streaming TextDecoder lab
Step 0 of 7Ready
Your turn: follow the blue line

A streaming decoder keeps unfinished multi-byte UTF-8 sequences between chunks. Step through a split é.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
const first = bytes.slice(0, 4);const second = bytes.slice(4); console.log(new TextDecoder("utf-8").decode(first)); const decoder = new TextDecoder("utf-8", { fatal: true });console.log(decoder.decode(first, { stream: true }));console.log(decoder.decode(second));
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

Base64 is byte text, not Unicode text

EDGE CASES

Base64 groups three bytes into four 6-bit indexes. Those indexes pick characters from an alphabet: A-Z, a-z, 0-9, +, and /. Padding = fills the last four-character group when the byte count is not a multiple of three. Base64url swaps + for -, / for _, and often omits padding.

The old browser helpers btoa and atob do not accept arbitrary JavaScript text. They operate on binary strings, where each string character's code is one byte from 0 through 255. That is why this first example works.

btoa and atob use binary stringsPop out in the code editor (opens in a new tab)JavaScript
console.log(btoa("é"));const binary = atob("6Q==");console.log(binary.charCodeAt(0));

The Euro sign is outside that one-byte range, so btoa("€") throws an InvalidCharacterError. The fix is not to guess a code page; encode the text to UTF-8 bytes first.

A deliberate btoa failurePop out in the code editor (opens in a new tab)JavaScript
console.log(btoa("€"));
Safe Unicode text to Base64 with TextEncoderPop out in the code editor (opens in a new tab)JavaScript
function bytesToBinaryString(bytes) {  return String.fromCharCode(...bytes);}function binaryStringToBytes(binary) {  return Uint8Array.from(binary, (character) => character.charCodeAt(0));} const bytes = new TextEncoder().encode("€");const base64 = btoa(bytesToBinaryString(bytes));console.log([...bytes].join(","));console.log(base64);console.log(new TextDecoder().decode(binaryStringToBytes(atob(base64))));
Base64 is reversible, not secret

Anyone can decode Base64 or base64url. Use it when a text-only place must carry bytes, such as JSON, URLs, or a data URL. Do not use it to hide API keys, tokens, or user data.

Modern Uint8Array helpers remove the binary-string detour

FEATURE DETECT

The typed-array Base64 and hex helpers are the direct APIs learners always wanted: bytes in, encoded text out; encoded text in, bytes out. MDN marks toBase64, fromBase64, toHex, and fromHex as standard-track, Baseline 2025 APIs with Chrome 140, Firefox 133, Safari 18.2, and Node 25 support. The CI runtime for this site is Node 22, so examples feature-detect and keep fallbacks.

Old and new byte encoders
APIShapeSupport note
btoa / atobBinary string onlyOld and widely available; every character must be one byte (0 through 255).
Uint8Array.toBase64()Bytes to Base64Standard-track typed-array helper. MDN marks Baseline 2025: Chrome 140, Firefox 133, Safari 18.2, Node 25.
Uint8Array.fromBase64()Base64 to bytesSupports alphabet: "base64url" and lastChunkHandling; feature-detect because Node 22 lacks it.
toHex() / fromHex()Bytes to hex and backSame Baseline 2025 browser support and Node 25+ in MDN data; use manual hex in Node 22.
Native Base64, base64url, and hex helpersPop out in the code editor (opens in a new tab)JavaScript
const bytes = new Uint8Array([104, 105, 63]);if (typeof bytes.toBase64 === "function" && typeof Uint8Array.fromBase64 === "function") {  console.log(bytes.toBase64());  console.log(bytes.toBase64({ alphabet: "base64url", omitPadding: true }));  console.log([...Uint8Array.fromBase64("aGk_", { alphabet: "base64url" })].join(","));  console.log([...Uint8Array.fromBase64("SGk", { lastChunkHandling: "loose" })].join(","));} else {  console.log("native Base64 helpers unavailable");} const token = new Uint8Array([222, 173, 190, 239]);if (typeof token.toHex === "function" && typeof Uint8Array.fromHex === "function") {  console.log(token.toHex());  console.log([...Uint8Array.fromHex("deadbeef")].join("-"));} else {  console.log([...token].map((byte) => byte.toString(16).padStart(2, "0")).join(""));}

The Base64 options are worth naming. { alphabet: 'base64url' } switches to URL-safe - and _. { omitPadding: true } drops trailing =. Uint8Array.fromBase64 also accepts lastChunkHandling; "loose" accepts unpadded final chunks, while stricter modes reject or stop before partial input.

Hex is simpler: one byte becomes exactly two base-16 digits. Until native toHex is available everywhere you run, this manual form is easy to audit.

Manual hex with padStartPop out in the code editor (opens in a new tab)JavaScript
const bytes = new Uint8Array([0, 15, 16, 255]);const hex = [...bytes]  .map((byte) => byte.toString(16).padStart(2, "0"))  .join("");console.log(hex); const roundTrip = Uint8Array.from(hex.match(/../g), (pair) => parseInt(pair, 16));console.log([...roundTrip].join(","));

Recipes you will use on real sites

PLAYGROUND

Use this workbench as a ruler. Type text and compare the UTF-16 code-unit count with code points, UTF-8 bytes, hex, Base64, and base64url. The visible source is intentionally small: encode once, then derive every byte-oriented representation from the same bytes.

Text-to-bytes workbench
Workbench sourcePop out in the code editor (opens in a new tab)JavaScript
const text = "Aé€😀";const bytes = new TextEncoder().encode(text);const hex = [...bytes].map((byte) => byte.toString(16).padStart(2, "0")).join("");const base64 = btoa(String.fromCharCode(...bytes)); console.log([...bytes].join(","));console.log(hex);console.log(base64);
Live encodings10 bytes
UTF-16 code units5
Unicode code points4
UTF-8 bytes65, 195, 169, 226, 130, 172, 240, 159, 152, 128
Hex41c3a9e282acf09f9880
Base64QcOp4oKs8J+YgA==
base64urlQcOp4oKs8J-YgA
Try it yourself

Every output starts from the same TextEncoder bytes. Change the text and compare code units, code points, bytes, hex, Base64, and base64url.

No learner code is evaluated. The playground only encodes the text you type with browser APIs.

JWTs demonstrate why “encoded” does not mean “private.” The middle part is base64url JSON. Decoding it is useful for display and debugging, but a security decision must verify the JWT signature on the server.

Decode a JWT payload partPop out in the code editor (opens in a new tab)JavaScript
function base64UrlToBytes(part) {  const base64 = part    .replace(/-/g, "+")    .replace(/_/g, "/")    .padEnd(Math.ceil(part.length / 4) * 4, "=");  const binary = atob(base64);  return Uint8Array.from(binary, (character) => character.charCodeAt(0));} const payloadPart = "eyJzdWIiOiIxMjMiLCJhZG1pbiI6ZmFsc2V9";const json = new TextDecoder().decode(base64UrlToBytes(payloadPart));console.log(json);console.log(JSON.parse(json).admin);

Data URLs are the opposite direction: keep small file bytes inside a string. They are handy for tiny generated files and demos, but large Base64 data URLs bloat HTML and CSS.

Build a small data URLPop out in the code editor (opens in a new tab)JavaScript
const bytes = new TextEncoder().encode("Hello");const base64 = btoa(String.fromCharCode(...bytes));console.log(`data:text/plain;base64,${base64}`);

Hashes use bytes too. The Web Crypto API is asynchronous, returns an ArrayBuffer, and leaves the display format to you. Hex is the usual digest display because it is stable, lowercase, and easy to compare.

SHA-256 digest as hexPop out in the code editor (opens in a new tab)JavaScript
async function sha256Hex(message) {  const bytes = new TextEncoder().encode(message);  const digest = await crypto.subtle.digest("SHA-256", bytes);  return [...new Uint8Array(digest)]    .map((byte) => byte.toString(16).padStart(2, "0"))    .join("");} sha256Hex("hello").then(console.log);
Base64 size overheadPop out in the code editor (opens in a new tab)JavaScript
for (const length of [1, 2, 3, 4, 12]) {  const base64Length = Math.ceil(length / 3) * 4;  console.log(length + " bytes -> " + base64Length + " Base64 chars");}
Choose the representation
  • new TextEncoder().encode(message)
  • new TextDecoder('utf-8', { fatal: true })
  • JWT payload part: eyJzdWIiOiIxMjMifQ
  • data:image/png;base64,...
  • 2cf24dba5fb0a30e...
  • byte.toString(16).padStart(2, '0')
Try it yourself
0 of 6 correct

Sort each real-world snippet by the conversion or byte display it needs.

Choose a category for every card. You can change an answer at any time; Reset clears them all.

Common misconceptions

  • “A JavaScript string is already UTF-8.” It is UTF-16 code units. UTF-8 appears after you encode it.
  • “btoa encodes Unicode text.” It encodes binary strings. Use TextEncoder for arbitrary text.
  • “Base64 makes data safe.” It only makes bytes text-safe. It is not encryption, signing, escaping, or validation.
  • “A decoder always throws on bad bytes.” The default inserts �. Add { fatal: true } when invalid input should stop.
  • “Hex and Base64 are interchangeable.” Both display bytes, but hex is two characters per byte while Base64 is denser and chunked.
Similar tools, different promises
Confusing ideaRealitySafer habit
btoa(text)Only safe for binary stringsEncode text with TextEncoder first, then Base64 the bytes.
Base64A transport encodingIt is reversible without a secret key, so it is not encryption.
HexA debugging-friendly encodingIt doubles size, but every two characters map to exactly one byte.
TextDecoder defaultReplacement modeBad bytes become �; add fatal: true when corrupt input should fail.

Practice exercises

5 EXERCISES
Exercise 1 · Warm-upPredict the UTF-8 bytes

Read the snippet and type the exact decimal bytes printed by the console.

Starter codePop out in the code editor (opens in a new tab)JavaScript
const bytes = new TextEncoder().encode("é");
console.log([...bytes].join(","));

Answer, then press Check. Spacing and letter case don’t matter.

    Exercise 2 · PracticeWrite the manual hex output

    Predict the hex string. This checks why padStart(2, "0") matters.

    Starter codePop out in the code editor (opens in a new tab)JavaScript
    const bytes = new Uint8Array([15, 16, 255]);
    const hex = [...bytes]
      .map((byte) => byte.toString(16).padStart(2, "0"))
      .join("");
    console.log(hex);

    Answer, then press Check. Spacing and letter case don’t matter.

      Exercise 3 · PracticeFind the Base64 bug

      The code throws. Type the browser API that should run before btoa for arbitrary Unicode text.

      Starter codePop out in the code editor (opens in a new tab)JavaScript
      console.log(btoa("€"));

      Answer, then press Check. Spacing and letter case don’t matter.

        Exercise 4 · PracticeRead a JWT payload

        Type the boolean printed by the final line.

        Starter codePop out in the code editor (opens in a new tab)JavaScript
        function base64UrlToBytes(part) {
          const base64 = part.replace(/-/g, "+").replace(/_/g, "/").padEnd(Math.ceil(part.length / 4) * 4, "=");
          return Uint8Array.from(atob(base64), (character) => character.charCodeAt(0));
        }
        const json = new TextDecoder().decode(base64UrlToBytes("eyJzdWIiOiIxMjMiLCJhZG1pbiI6ZmFsc2V9"));
        console.log(JSON.parse(json).admin);

        Answer, then press Check. Spacing and letter case don’t matter.

          Exercise 5 · ChallengeEstimate Base64 size

          A payload has 5 bytes before encoding. How many Base64 characters will represent it?

          Starter codePop out in the code editor (opens in a new tab)JavaScript
          const length = 5;
          console.log(Math.ceil(length / 3) * 4);

          Answer, then press Check. Spacing and letter case don’t matter.

            Check your understanding

            7 QUESTIONS
            Lesson quiz · 7 questionsScore: first tries count
            1. Question 1 of 7Which statement best describes JavaScript strings versus files or network data?

              Choose an answer to see the explanation.

            2. Question 2 of 7What bytes does UTF-8 use for €?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              const bytes = new TextEncoder().encode("€");
              console.log([...bytes].join(","));

              Choose an answer to see the explanation.

            3. Question 3 of 7What does replacement-mode decoding print for invalid UTF-8?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              const bad = new Uint8Array([0xC3, 0x28]);
              console.log(new TextDecoder("utf-8").decode(bad));

              Choose an answer to see the explanation.

            4. Question 4 of 7What does btoa('é') print in a browser?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              console.log(btoa("é"));

              Choose an answer to see the explanation.

            5. Question 5 of 7Which Base64 form is URL-safe and omits padding?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              const base64 = "aGk/";
              console.log(base64.replace(/\+/g, "-").replace(/\//g, "_").replace(/=+$/g, ""));

              Choose an answer to see the explanation.

            6. Question 6 of 7What hex string does the manual encoder produce?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              const bytes = new Uint8Array([0, 15, 255]);
              console.log([...bytes].map((byte) => byte.toString(16).padStart(2, "0")).join(""));

              Choose an answer to see the explanation.

            7. Question 7 of 7Why is Base64 about one third larger than the original bytes?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              for (const length of [3, 6]) {
                console.log(Math.ceil(length / 3) * 4);
              }

              Choose an answer to see the explanation.

            Key takeaways

            • JavaScript strings are UTF-16 code units; most storage and transport APIs work with bytes.
            • TextEncoder always writes UTF-8; TextDecoder reads bytes with labels and options.
            • UTF-8 uses one to four bytes per code point; byte length is not string length.
            • btoa and atob use binary strings, so encode Unicode text to bytes first.
            • Base64 and hex represent bytes as text. They are reversible encodings, not encryption.
            • Native Uint8Array Base64 and hex helpers are useful, but feature-detect for Node 22 and older browsers.

            Remember the one-liner.
            Encode text before treating it as bytes, and label byte display formats like Base64 and hex for what they are: transport and debugging tools, not secrecy.

            Up next: memory management and garbage collection, where the lifetime of objects, buffers, and references becomes the main story.

            CompleteFrontend Clear concepts. Working examples.