Unicode & string internals
Learn why an emoji has length 2, how JavaScript strings store UTF-16 code units, how to count code points and graphemes, and how normalization and well-formed text keep user input safe.
- 01Name the three countsTell code units, code points, and graphemes apart when reading string results.
- 02Use Unicode APIs safelyChoose
codePointAt,fromCodePoint,normalize, andIntl.Segmenterfor the right job. - 03Protect text boundariesAvoid broken emoji, normalize comparisons, and repair or reject ill-formed strings.
Why an emoji has length 2
JavaScript strings look simple: quotes around text. Under the hood, though, every string is a sequence of UTF-16 code units. That is why the published Strings lesson could tease that "😀".length is 2. A person sees one smiling face. JavaScript sees two 16-bit storage boxes that work together.
This lesson is the X-ray view. You will separate three ideas that are often called “characters” in everyday speech: the storage boxes JavaScript indexes, the Unicode numbers those boxes encode, and the human-visible graphemes users expect your UI to respect.
Unicode gives text a giant catalog of numbers called code points. JavaScript stores those numbers as UTF-16 code units. Users usually care about graphemes: what they see as one character.
Imagine a worldwide catalog. Every item has a stable number. A local office still needs a paper form to store that number, and the form has fixed-width boxes. Short numbers fit in one box; long numbers need two boxes side by side. Unicode is the catalog. UTF-16 is the form JavaScript strings use.
- In real life: A catalog assigns every product a number
- In JavaScript: Unicode assigns every character a code point like U+1F600
- In real life: Most names fit on one form line
- In JavaScript: Most code points fit in one UTF-16 code unit
- In real life: A long name spills onto two lines
- In JavaScript: Emoji and many rare characters need a surrogate pair
- In real life: A customer sees one product
- In JavaScript: A user sees one character even when storage uses two boxes
Where the analogy stops: The form analogy explains storage size, not language rules. Human writing systems have their own rules, which is why grapheme segmentation exists.
We will mention two deeper future lessons by title only: Text encoding & Base64 explains UTF-8 bytes and TextEncoder, and Strings inside the engine explains engine storage tricks. Today stays at the language level you use in application code.
UTF-16 code units: the boxes JavaScript indexes
INTERACTIVEA code unit is one 16-bit box in the string. The APIs you already know from String methods often use these boxes directly: length, bracket indexes, charCodeAt, and slice. That is useful for old APIs and low-level inspection, but it is not the same as “what the user sees.”
Use the X-ray. Try héllo, 😀, 👍🏽, a family emoji, composed and decomposed é, and a flag. Watch how the three counts disagree.
const text = "😀";text.length;text.split("");[...text];Array.from(new Intl.Segmenter(undefined, { granularity: "grapheme" }).segment(text));split("")["\ud83d","\ude00"]Code units
0 0xD83D
1 0xDE00
Code points
0 😀 U+1F600
Graphemes
0 😀
This browser reports 2 code units, 1 code points, and 1 grapheme. Intl.Segmenter is being checked.
Notice the most surprising column: split("") splits by code unit, so it can tear an emoji into halves. Spread syntax, [...text], uses string iteration, so it walks code points instead. The published for…of & for…in lesson uses the same iteration behavior.
| Unit | How JavaScript counts it | Example |
|---|---|---|
| Code unit | One UTF-16 16-bit box; length, indexes, charCodeAt, and slice use these | "😀".length is 2 |
| Code point | One Unicode catalog number; for…of, spread, codePointAt, and fromCodePoint understand these | [..."😀"].length is 1 |
| Grapheme | One user-perceived character; Intl.Segmenter with granularity: "grapheme" segments these | 👍🏽 is one grapheme but two code points |
Code points & surrogate pairs
A code point is Unicode’s catalog number for a character-like thing. The smiling face is U+1F600. The letter A is U+0041. Many code points fit in one UTF-16 code unit, but code points above U+FFFF need two code units. That pair is called a surrogate pair: a high surrogate followed by a low surrogate.
If a venue prints a ticket across two detachable halves, the whole ticket works only when both halves stay together. A surrogate pair works the same way. JavaScript can hold one half, but that half is broken text.
- In real life: One concert ticket can have a left and right half
- In JavaScript: One code point can be stored as a high and low surrogate
- In real life: Either half alone will not get you in
- In JavaScript: A lone surrogate is not a valid character
- In real life: The scanner reads both halves together
- In JavaScript:
codePointAt(0)reads the pair when it starts at the high half
Where the analogy stops: Surrogate pairs are mechanical UTF-16 encoding pieces. They do not always match a user-visible character, because some graphemes contain multiple code points.
Step through a real emoji. Every number below is produced by JavaScript: 55357 is 0xD83D, the first code unit; 128512 is 0x1F600, the code point.
Predict each result before stepping. This replay is about the gap between what a person sees and what UTF-16 stores.
script
console.log(smile.length);console.log(smile.charCodeAt(0));console.log(smile.codePointAt(0));console.log(String.fromCodePoint(128512));console.log(String.fromCharCode(0xD83D, 0xDE00));console.log(smile.slice(0, 1));codePointAt, fromCodePoint & fromCharCode
STEP THROUGHUse charCodeAt and fromCharCode when you deliberately want UTF-16 code units. Use codePointAt and String.fromCodePoint when you mean Unicode code points. The names are literal: “char code” is the older code-unit view; “code point” is the Unicode catalog view.
codePointAt(index) only combines a surrogate pair when index points at the high surrogate. If you pass the low surrogate index, it returns that low code unit by itself. Iterating with for…of or spread avoids landing in the middle.
| API | Unit | Use it when |
|---|---|---|
charCodeAt(index) | Code unit | You need the exact UTF-16 box at an index |
codePointAt(index) | Code point | You want the Unicode number starting at this position |
String.fromCharCode(...) | Code units | You already have UTF-16 boxes, including both surrogate halves |
String.fromCodePoint(...) | Code points | You have Unicode catalog numbers such as 0x1F600 |
The next sorter makes you choose the counting level. Do not worry if you miss one; each card explains the boundary it uses.
str.lengthstr.charCodeAt(0)str.slice(0, 1)[...str]str.codePointAt(0)String.fromCodePoint(0x1F600)Intl.Segmenter(..., { granularity: "grapheme" })- A visible character limit in a profile name
Sort each operation by the unit it counts or protects.
normalize(): one spelling for the same word
STEP THROUGHUnicode sometimes has more than one valid spelling for the same visible text. The letter é can be stored as one prebuilt code point, U+00E9, or as e followed by a combining acute accent, U+0301. They look identical in most fonts, but strict equality compares stored code units, not appearance.
If one list says “A. Lovelace” and another says “Ada Lovelace,” a human sees the connection. A computer needs a rule. Normalization is that rule for Unicode spelling variants.
- In real life: Two people write the same name with different abbreviations
- In JavaScript: Two strings look the same but store different code points
- In real life: A form asks everyone to use the official spelling
- In JavaScript:
normalize()rewrites text to a canonical form - In real life: Then the names can be compared fairly
- In JavaScript: Normalized strings can be compared, searched, or deduplicated
Where the analogy stops: Normalization is not translation, spell-checking, or case folding. It only changes canonical Unicode forms.
The two strings look identical. Step through to see why raw equality and normalized equality disagree.
script
const decomposed = "é";console.log(composed === decomposed);console.log(composed.length);console.log(decomposed.length);console.log(composed.normalize() === decomposed.normalize());console.log(composed.normalize("NFD").length);Practical rule: normalize before comparing user input, storing keys, deduplicating names, or doing accent-sensitive search. The default normalize() uses NFC, the composed form. NFD splits many prebuilt characters into base marks and combining marks, which can be useful for accent stripping or teaching, but NFC is the everyday comparison choice.
Graphemes with Intl.Segmenter
INTERACTIVEA grapheme is what a user usually experiences as one character. It can be built from several code points glued together: a thumbs-up plus a skin-tone modifier, a family joined with zero-width joiners, a flag made of regional indicator symbols, or an e plus a combining accent.
Intl.Segmenter is available in Node 16+ and modern browsers. This page checks for it before using it. If you support an older browser, load a tested grapheme-splitting library or fall back to a less exact count with a clear product decision.
The common bug is truncating with slice. That cuts by code units, so it can show a broken replacement box or remove a skin tone from an emoji. The safe version segments first, then slices the list of graphemes.
function firstGraphemes(text, limit) { const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" }); return Array.from(segmenter.segment(text), part => part.segment) .slice(0, limit) .join("");} console.log("👍🏽 launch".slice(0, 2));console.log(firstGraphemes("👍🏽 launch", 1));Code-unit slice gives "\ud83d". Segmenter-based truncation gives "👍🏽" and keeps whole graphemes.
slice, so it cuts by UTF-16 code units. The safe version uses grapheme segments.isWellFormed & toWellFormed
STEP THROUGHA JavaScript string can contain a lone surrogate: one half of a surrogate pair with the other half missing. That string is ill-formed. Recent browsers and Node 20+ provide isWellFormed() to detect it and toWellFormed() to replace broken halves with U+FFFD, the replacement character.
Why should you care? Encoders that send text out of JavaScript often require valid Unicode scalar values. For example, encodeURIComponent("\uD83D") throws a URIError. Repair, reject, or re-request broken text before sending it to servers or URLs.
Step through an intentionally broken string. Recent browsers and Node 20+ have isWellFormed and toWellFormed; the page feature-checks them.
script
console.log(broken.isWellFormed());console.log(broken.toWellFormed());try { encodeURIComponent(broken);} catch (error) { console.log(error.name);}The lesson’s examples feature-check these methods in the site. If your runtime lacks them, you can still validate surrogate pairs with a small helper or a well-maintained library. Prefer feature checks over assuming every visitor has the newest engine.
Where you’ll use this
Unicode internals show up whenever text crosses a boundary: a profile name, a search box, a slug generator, a database key, or a UI counter. Most bugs come from using a lower-level count for a higher-level job.
| Job | Good choice | Why |
|---|---|---|
| Inspect a protocol or legacy API that talks UTF-16 | Code units | The external API is defined in UTF-16 boxes |
| Reverse simple emoji text for a kata | Code points with spread | Keeps surrogate pairs together, but not every grapheme cluster |
| Limit a display name to 20 visible characters | Graphemes | Users expect 👍🏽 and 🇯🇵 to count as one visible character each |
| Compare user input for search or dedupe | Normalize first | Canonical spellings like é and e + accent should match |
| Encode text into a URL | Check well-formed text | Lone surrogates make encodeURIComponent throw |
The next lesson, Comparing & sorting text, builds on this: normalize before comparing when canonical forms matter, and use localeCompare or Intl.Collator when a user’s language decides order.
Common misconceptions
“length counts characters.”
It counts UTF-16 code units. That matches many older English-only examples, but not emoji, combining marks, or many rare characters.
“charCodeAt gives the Unicode character number.”
It gives one code unit. Use codePointAt when you mean a code point and might meet surrogate pairs.
“Spread always counts user-visible characters.”
Spread counts code points. It keeps 😀 together, but 👍🏽 is still two code points and one grapheme.
“If two strings look the same, === is true.”
Strict equality compares stored sequences. Normalize first when canonical Unicode spellings should be treated as the same text.
“JavaScript strings are UTF-8.”
JavaScript string indexing is defined in UTF-16 code units. UTF-8 bytes matter when encoding text for files, networks, and the later Text encoding & Base64 lesson.
Practice: Unicode strings
5 EXERCISESPredict all four outputs before running the snippet.
console.log("😀".length);
console.log([..."😀"].length);
console.log("👍🏽".length);
console.log([..."👍🏽"].length);The outputs are 2, 1, 4, and 2. The thumbs-up with skin tone is two code points and four UTF-16 code units.
Reverse Hi 😀 in a way that does not split the emoji.
const text = "Hi 😀";
console.log([...text].reverse().join(""));const text = "Hi 😀";
console.log([...text].reverse().join(""));[...text] iterates by code point. That is enough for this string, so the emoji remains whole and the output is 😀 iH. Full grapheme-safe reversing is trickier for families and skin tones.
What do the two comparisons print?
const typed = "Café";
const saved = "Café";
console.log(typed === saved);
console.log(typed.normalize() === saved.normalize());The raw strings differ, so the first output is false. After default NFC normalization, both use the same canonical spelling, so the second output is true.
How many visible characters should a profile-name counter show for 👍🏽👍?
const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" });
const text = "👍🏽👍";
console.log(Array.from(segmenter.segment(text)).length);The string 👍🏽👍 has two graphemes: 👍🏽 and 👍. It has more code units than visible characters, but the UI limit should count 2.
Show the first visible character from 👍🏽 launch without breaking it.
const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" });
const text = "👍🏽 launch";
const safe = Array.from(segmenter.segment(text), part => part.segment).slice(0, 1).join("");
console.log(safe);const segmenter = new Intl.Segmenter(undefined, { granularity: "grapheme" });
const text = "👍🏽 launch";
const safe = Array.from(segmenter.segment(text), part => part.segment).slice(0, 1).join("");
console.log(safe);The code builds grapheme segments, takes the first one, and joins it back. The result is the whole 👍🏽 grapheme, not a broken surrogate or a missing skin tone.
Quiz: check your understanding
7 QUESTIONSEvery answer explains itself. Read the wrong answers too; Unicode mistakes usually come from picking the wrong unit.
Question 1 of 7What does
String.prototype.lengthcount?Choose an answer to see the explanation.
Question 2 of 7What does the emoji length snippet print?
Read the code, then predictconsole.log("😀".length);Choose an answer to see the explanation.
Question 3 of 7What does the emoji code point snippet print?
Read the code, then predictconsole.log("😀".codePointAt(0));Choose an answer to see the explanation.
Question 4 of 7Why normalize before comparing some user input?
Choose an answer to see the explanation.
Question 5 of 7What do the normalization comparisons print?
Read the code, then predictconsole.log("é" === "e\u0301"); console.log("é".normalize() === "e\u0301".normalize());Choose an answer to see the explanation.
Question 6 of 7Which tool should you use to count visible characters for a UI limit?
Choose an answer to see the explanation.
Question 7 of 7What happens with a lone surrogate in
encodeURIComponent?Read the code, then predicttry { encodeURIComponent(String.fromCharCode(0xD83D)); } catch (error) { console.log(error.name); }Choose an answer to see the explanation.
Key takeaways
- JavaScript strings are indexed as UTF-16 code units.
- Code points are Unicode catalog numbers; surrogate pairs store code points above U+FFFF.
codePointAtandString.fromCodePointare the code-point APIs;charCodeAtandfromCharCodeare code-unit APIs.normalize()makes canonically equivalent spellings comparable.Intl.Segmentercounts graphemes for UI limits and safe truncation.isWellFormedandtoWellFormedhelp detect and repair lone surrogates before encoding.
Remember the one-liner.
Code units are storage boxes, code points are Unicode numbers, and graphemes are what users see.
Up next: Comparing & sorting text, where normalization meets localeCompare and language-aware ordering.