cf.completefrontendCode editorOpen lab
THE JAVASCRIPT FIELD GUIDE

Unicode & string internals

Learn why an emoji has length 2, how JavaScript strings store UTF-16 code units, how to count code points and graphemes, and how normalization and well-formed text keep user input safe.

By the end, you can
  • 01
    Name the three countsTell code units, code points, and graphemes apart when reading string results.
  • 02
    Use Unicode APIs safelyChoose codePointAt, fromCodePoint, normalize, and Intl.Segmenter for the right job.
  • 03
    Protect text boundariesAvoid broken emoji, normalize comparisons, and repair or reject ill-formed strings.

Why an emoji has length 2

JavaScript strings look simple: quotes around text. Under the hood, though, every string is a sequence of UTF-16 code units. That is why the published Strings lesson could tease that "😀".length is 2. A person sees one smiling face. JavaScript sees two 16-bit storage boxes that work together.

This lesson is the X-ray view. You will separate three ideas that are often called “characters” in everyday speech: the storage boxes JavaScript indexes, the Unicode numbers those boxes encode, and the human-visible graphemes users expect your UI to respect.

The short version

Unicode gives text a giant catalog of numbers called code points. JavaScript stores those numbers as UTF-16 code units. Users usually care about graphemes: what they see as one character.

Real-life analogyUnicode is a catalog; UTF-16 is the form

Imagine a worldwide catalog. Every item has a stable number. A local office still needs a paper form to store that number, and the form has fixed-width boxes. Short numbers fit in one box; long numbers need two boxes side by side. Unicode is the catalog. UTF-16 is the form JavaScript strings use.

In real life: A catalog assigns every product a number
In JavaScript: Unicode assigns every character a code point like U+1F600
In real life: Most names fit on one form line
In JavaScript: Most code points fit in one UTF-16 code unit
In real life: A long name spills onto two lines
In JavaScript: Emoji and many rare characters need a surrogate pair
In real life: A customer sees one product
In JavaScript: A user sees one character even when storage uses two boxes

Where the analogy stops: The form analogy explains storage size, not language rules. Human writing systems have their own rules, which is why grapheme segmentation exists.

We will mention two deeper future lessons by title only: Text encoding & Base64 explains UTF-8 bytes and TextEncoder, and Strings inside the engine explains engine storage tricks. Today stays at the language level you use in application code.

UTF-16 code units: the boxes JavaScript indexes

INTERACTIVE

A code unit is one 16-bit box in the string. The APIs you already know from String methods often use these boxes directly: length, bracket indexes, charCodeAt, and slice. That is useful for old APIs and low-level inspection, but it is not the same as “what the user sees.”

Use the X-ray. Try héllo, 😀, 👍🏽, a family emoji, composed and decomposed é, and a flag. Watch how the three counts disagree.

String X-ray playground
Operations the X-ray runsPop out in the code editor (opens in a new tab)JavaScript
const text = "😀";text.length;text.split("");[...text];Array.from(new Intl.Segmenter(undefined, { granularity: "grapheme" }).segment(text));
Computed nowChecking
length (code units)2
[...text].length (code points)1
grapheme count1
split("")["\ud83d","\ude00"]
[...text]["😀"]

Code units

0 0xD83D

1 0xDE00

Code points

0 😀 U+1F600

Graphemes

0 😀

Try it yourself

This browser reports 2 code units, 1 code points, and 1 grapheme. Intl.Segmenter is being checked.

All values are computed in your browser. The playground feature-checks Intl.Segmenter and shows a fallback if needed.

Notice the most surprising column: split("") splits by code unit, so it can tear an emoji into halves. Spread syntax, [...text], uses string iteration, so it walks code points instead. The published for…of & for…in lesson uses the same iteration behavior.

The three text units you must keep separate
UnitHow JavaScript counts itExample
Code unitOne UTF-16 16-bit box; length, indexes, charCodeAt, and slice use these"😀".length is 2
Code pointOne Unicode catalog number; for…of, spread, codePointAt, and fromCodePoint understand these[..."😀"].length is 1
GraphemeOne user-perceived character; Intl.Segmenter with granularity: "grapheme" segments these👍🏽 is one grapheme but two code points

Code points & surrogate pairs

A code point is Unicode’s catalog number for a character-like thing. The smiling face is U+1F600. The letter A is U+0041. Many code points fit in one UTF-16 code unit, but code points above U+FFFF need two code units. That pair is called a surrogate pair: a high surrogate followed by a low surrogate.

Real-life analogyA surrogate pair is a two-part ticket

If a venue prints a ticket across two detachable halves, the whole ticket works only when both halves stay together. A surrogate pair works the same way. JavaScript can hold one half, but that half is broken text.

In real life: One concert ticket can have a left and right half
In JavaScript: One code point can be stored as a high and low surrogate
In real life: Either half alone will not get you in
In JavaScript: A lone surrogate is not a valid character
In real life: The scanner reads both halves together
In JavaScript: codePointAt(0) reads the pair when it starts at the high half

Where the analogy stops: Surrogate pairs are mechanical UTF-16 encoding pieces. They do not always match a user-visible character, because some graphemes contain multiple code points.

Step through a real emoji. Every number below is produced by JavaScript: 55357 is 0xD83D, the first code unit; 128512 is 0x1F600, the code point.

X-ray one emoji
Step 0 of 7Ready
Your turn: follow the blue line

Predict each result before stepping. This replay is about the gap between what a person sees and what UTF-16 stores.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
console.log(smile.length);console.log(smile.charCodeAt(0));console.log(smile.codePointAt(0));console.log(String.fromCodePoint(128512));console.log(String.fromCharCode(0xD83D, 0xDE00));console.log(smile.slice(0, 1));
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

codePointAt, fromCodePoint & fromCharCode

STEP THROUGH

Use charCodeAt and fromCharCode when you deliberately want UTF-16 code units. Use codePointAt and String.fromCodePoint when you mean Unicode code points. The names are literal: “char code” is the older code-unit view; “code point” is the Unicode catalog view.

Important edge case

codePointAt(index) only combines a surrogate pair when index points at the high surrogate. If you pass the low surrogate index, it returns that low code unit by itself. Iterating with for…of or spread avoids landing in the middle.

Which API means which unit?
APIUnitUse it when
charCodeAt(index)Code unitYou need the exact UTF-16 box at an index
codePointAt(index)Code pointYou want the Unicode number starting at this position
String.fromCharCode(...)Code unitsYou already have UTF-16 boxes, including both surrogate halves
String.fromCodePoint(...)Code pointsYou have Unicode catalog numbers such as 0x1F600

The next sorter makes you choose the counting level. Do not worry if you miss one; each card explains the boundary it uses.

Which count do you need?
  • str.length
  • str.charCodeAt(0)
  • str.slice(0, 1)
  • [...str]
  • str.codePointAt(0)
  • String.fromCodePoint(0x1F600)
  • Intl.Segmenter(..., { granularity: "grapheme" })
  • A visible character limit in a profile name
Try it yourself
0 of 8 correct

Sort each operation by the unit it counts or protects.

Choose a category for every card. You can change an answer at any time; Reset clears them all.

normalize(): one spelling for the same word

STEP THROUGH

Unicode sometimes has more than one valid spelling for the same visible text. The letter é can be stored as one prebuilt code point, U+00E9, or as e followed by a combining acute accent, U+0301. They look identical in most fonts, but strict equality compares stored code units, not appearance.

Real-life analogyNormalization is agreeing on one spelling

If one list says “A. Lovelace” and another says “Ada Lovelace,” a human sees the connection. A computer needs a rule. Normalization is that rule for Unicode spelling variants.

In real life: Two people write the same name with different abbreviations
In JavaScript: Two strings look the same but store different code points
In real life: A form asks everyone to use the official spelling
In JavaScript: normalize() rewrites text to a canonical form
In real life: Then the names can be compared fairly
In JavaScript: Normalized strings can be compared, searched, or deduplicated

Where the analogy stops: Normalization is not translation, spell-checking, or case folding. It only changes canonical Unicode forms.

Normalize two identical-looking strings
Step 0 of 7Ready
Your turn: follow the blue line

The two strings look identical. Step through to see why raw equality and normalized equality disagree.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
const decomposed = "é";console.log(composed === decomposed);console.log(composed.length);console.log(decomposed.length);console.log(composed.normalize() === decomposed.normalize());console.log(composed.normalize("NFD").length);
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

Practical rule: normalize before comparing user input, storing keys, deduplicating names, or doing accent-sensitive search. The default normalize() uses NFC, the composed form. NFD splits many prebuilt characters into base marks and combining marks, which can be useful for accent stripping or teaching, but NFC is the everyday comparison choice.

Graphemes with Intl.Segmenter

INTERACTIVE

A grapheme is what a user usually experiences as one character. It can be built from several code points glued together: a thumbs-up plus a skin-tone modifier, a family joined with zero-width joiners, a flag made of regional indicator symbols, or an e plus a combining accent.

Feature check for real apps

Intl.Segmenter is available in Node 16+ and modern browsers. This page checks for it before using it. If you support an older browser, load a tested grapheme-splitting library or fall back to a less exact count with a clear product decision.

The common bug is truncating with slice. That cuts by code units, so it can show a broken replacement box or remove a skin tone from an emoji. The safe version segments first, then slices the list of graphemes.

Truncate without breaking emoji
Safe grapheme truncationPop out in the code editor (opens in a new tab)JavaScript
function firstGraphemes(text, limit) {  const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" });  return Array.from(segmenter.segment(text), part => part.segment)    .slice(0, limit)    .join("");} console.log("👍🏽 launch".slice(0, 2));console.log(firstGraphemes("👍🏽 launch", 1));
Result1 wanted
text.slice(0, 1)"\ud83d"
firstGraphemes(text, 1)"👍🏽"
browser has isWellFormedchecking
Try it yourself

Code-unit slice gives "\ud83d". Segmenter-based truncation gives "👍🏽" and keeps whole graphemes.

The unsafe version uses slice, so it cuts by UTF-16 code units. The safe version uses grapheme segments.

isWellFormed & toWellFormed

STEP THROUGH

A JavaScript string can contain a lone surrogate: one half of a surrogate pair with the other half missing. That string is ill-formed. Recent browsers and Node 20+ provide isWellFormed() to detect it and toWellFormed() to replace broken halves with U+FFFD, the replacement character.

Why should you care? Encoders that send text out of JavaScript often require valid Unicode scalar values. For example, encodeURIComponent("\uD83D") throws a URIError. Repair, reject, or re-request broken text before sending it to servers or URLs.

Detect and repair a lone surrogate
Step 0 of 5Ready
Your turn: follow the blue line

Step through an intentionally broken string. Recent browsers and Node 20+ have isWellFormed and toWellFormed; the page feature-checks them.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
console.log(broken.isWellFormed());console.log(broken.toWellFormed());try {  encodeURIComponent(broken);} catch (error) {  console.log(error.name);}
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.
Compatibility note

The lesson’s examples feature-check these methods in the site. If your runtime lacks them, you can still validate surrogate pairs with a small helper or a well-maintained library. Prefer feature checks over assuming every visitor has the newest engine.

Where you’ll use this

Unicode internals show up whenever text crosses a boundary: a profile name, a search box, a slug generator, a database key, or a UI counter. Most bugs come from using a lower-level count for a higher-level job.

Pick the count that matches the job
JobGood choiceWhy
Inspect a protocol or legacy API that talks UTF-16Code unitsThe external API is defined in UTF-16 boxes
Reverse simple emoji text for a kataCode points with spreadKeeps surrogate pairs together, but not every grapheme cluster
Limit a display name to 20 visible charactersGraphemesUsers expect 👍🏽 and 🇯🇵 to count as one visible character each
Compare user input for search or dedupeNormalize firstCanonical spellings like é and e + accent should match
Encode text into a URLCheck well-formed textLone surrogates make encodeURIComponent throw
Tease: comparing and sorting text

The next lesson, Comparing & sorting text, builds on this: normalize before comparing when canonical forms matter, and use localeCompare or Intl.Collator when a user’s language decides order.

Common misconceptions

“length counts characters.”

It counts UTF-16 code units. That matches many older English-only examples, but not emoji, combining marks, or many rare characters.

“charCodeAt gives the Unicode character number.”

It gives one code unit. Use codePointAt when you mean a code point and might meet surrogate pairs.

“Spread always counts user-visible characters.”

Spread counts code points. It keeps 😀 together, but 👍🏽 is still two code points and one grapheme.

“If two strings look the same, === is true.”

Strict equality compares stored sequences. Normalize first when canonical Unicode spellings should be treated as the same text.

“JavaScript strings are UTF-8.”

JavaScript string indexing is defined in UTF-16 code units. UTF-8 bytes matter when encoding text for files, networks, and the later Text encoding & Base64 lesson.

Practice: Unicode strings

5 EXERCISES
Exercise 1 · Warm-upPredict the counts

Predict all four outputs before running the snippet.

Starter codePop out in the code editor (opens in a new tab)JavaScript
console.log("😀".length);
console.log([..."😀"].length);
console.log("👍🏽".length);
console.log([..."👍🏽"].length);

Answer, then press Check. Spacing and letter case don’t matter.

    Exercise 2 · PracticeReverse without splitting the smile

    Reverse Hi 😀 in a way that does not split the emoji.

    Starter codePop out in the code editor (opens in a new tab)JavaScript
    const text = "Hi 😀";
    console.log([...text].reverse().join(""));

    Answer, then press Check. Spacing and letter case don’t matter.

      Exercise 3 · PracticeCompare café safely

      What do the two comparisons print?

      Starter codePop out in the code editor (opens in a new tab)JavaScript
      const typed = "Café";
      const saved = "Café";
      console.log(typed === saved);
      console.log(typed.normalize() === saved.normalize());

      Answer, then press Check. Spacing and letter case don’t matter.

        Exercise 4 · PracticeCount a character limit

        How many visible characters should a profile-name counter show for 👍🏽👍?

        Starter codePop out in the code editor (opens in a new tab)JavaScript
        const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" });
        const text = "👍🏽👍";
        console.log(Array.from(segmenter.segment(text)).length);

        Answer, then press Check. Spacing and letter case don’t matter.

          Exercise 5 · ChallengeFix a broken truncation

          Show the first visible character from 👍🏽 launch without breaking it.

          Starter codePop out in the code editor (opens in a new tab)JavaScript
          const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" });
          const text = "👍🏽 launch";
          const safe = Array.from(segmenter.segment(text), part => part.segment).slice(0, 1).join("");
          console.log(safe);

          Answer, then press Check. Spacing and letter case don’t matter.

            Quiz: check your understanding

            7 QUESTIONS

            Every answer explains itself. Read the wrong answers too; Unicode mistakes usually come from picking the wrong unit.

            Lesson quiz · 7 questionsScore: first tries count
            1. Question 1 of 7What does String.prototype.length count?

              Choose an answer to see the explanation.

            2. Question 2 of 7What does the emoji length snippet print?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              console.log("😀".length);

              Choose an answer to see the explanation.

            3. Question 3 of 7What does the emoji code point snippet print?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              console.log("😀".codePointAt(0));

              Choose an answer to see the explanation.

            4. Question 4 of 7Why normalize before comparing some user input?

              Choose an answer to see the explanation.

            5. Question 5 of 7What do the normalization comparisons print?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              console.log("é" === "e\u0301");
              console.log("é".normalize() === "e\u0301".normalize());

              Choose an answer to see the explanation.

            6. Question 6 of 7Which tool should you use to count visible characters for a UI limit?

              Choose an answer to see the explanation.

            7. Question 7 of 7What happens with a lone surrogate in encodeURIComponent?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              try {
                encodeURIComponent(String.fromCharCode(0xD83D));
              } catch (error) {
                console.log(error.name);
              }

              Choose an answer to see the explanation.

            Key takeaways

            • JavaScript strings are indexed as UTF-16 code units.
            • Code points are Unicode catalog numbers; surrogate pairs store code points above U+FFFF.
            • codePointAt and String.fromCodePoint are the code-point APIs; charCodeAt and fromCharCode are code-unit APIs.
            • normalize() makes canonically equivalent spellings comparable.
            • Intl.Segmenter counts graphemes for UI limits and safe truncation.
            • isWellFormed and toWellFormed help detect and repair lone surrogates before encoding.

            Remember the one-liner.
            Code units are storage boxes, code points are Unicode numbers, and graphemes are what users see.

            Up next: Comparing & sorting text, where normalization meets localeCompare and language-aware ordering.

            CompleteFrontend Clear concepts. Working examples.