Comparing & sorting text
Compare strings by code units, sort human text with localeCompare and Intl.Collator, handle locale-aware case conversion, and build case-insensitive search that respects accents.
- 01Explain default orderPredict when code-unit comparison puts capitals or accents in surprising places.
- 02Choose locale toolsUse
localeComparefor one-off comparisons and reuseIntl.Collatorfor bigger sorts. - 03Search kindlyFold case and accents deliberately instead of assuming
toLowerCase()solves every language.
Humans and code units want different things
JavaScript can compare any two strings immediately, but its fastest default answer is not a human dictionary. It compares the numbers inside the string: UTF-16 code units. You met those numbers in Unicode & string internals. They are perfect for a precise machine rule and surprising for names, titles, product lists, and search boxes.
This lesson is the bridge between simple Comparisons, array Sorting & copying safely, and the later Internationalization with Intl module. You will still write ordinary comparators, but the comparison itself will come from text-aware tools.
Text collation is the set of rules that decides whether one string should come before, after, or equal to another string for a particular language and purpose.
A machine can sort names by the codes behind their letters. Your phone sorts contacts in an order people expect for a chosen language. JavaScript can do either job.
- In real life: The code behind each letter
- In JavaScript: A character’s UTF-16 code unit
- In real life: Sorting those codes
- In JavaScript: Default
<,>, and stringsort() - In real life: Your phone’s contact order
- In JavaScript: Calling
localeComparewith a locale - In real life: Keeping one contact-list setting
- In JavaScript: Reusing one
Intl.Collator
Where the analogy stops: A phone chooses its own contact rules. JavaScript follows the locale and options you provide.
Default string comparison is code-unit comparison
STEP THROUGHWhen both sides of < are strings, JavaScript compares the first code unit where they differ. The same default appears when Array.prototype.sort() sorts strings without a comparator.
console.log("a".charCodeAt(0)); // 97console.log("Z".charCodeAt(0)); // 90 console.log("Zebra" < "apple"); // trueCode-unit order is not wrong. It is just a machine promise, not a language promise. It also explains why "10" < "9": the first code unit "1" comes before "9".
Switch the comparison method, predict the printed order, then step through the recorded run.
script
const sorted = [...words].sort();console.log(sorted.join(", "));Use default order for IDs, exact protocol strings, and places where you truly want code-unit order. For names and labels shown to users, choose a locale-aware comparison.
localeCompare asks for dictionary-like order
INTERACTIVEEvery string has localeCompare. It returns a negative number, zero, or a positive number: before, equal for this comparison, or after. That shape is exactly what sort comparators need.
"äpfel".localeCompare("zebra", "de"); // negative"äpfel".localeCompare("zebra", "sv"); // positive[...names].sort((a, b) => a.localeCompare(b, "de"));The important habit is to pass an explicit locale such as "en", "de", "sv", or "tr". If you omit it, JavaScript uses the environment’s default locale. That can be useful for an app configured from the user’s language, but examples and tests should be explicit.
const words = ["zebra", "äpfel", "apfel", "åre", "öl"];const collator = new Intl.Collator(locale);const sorted = [...words].sort(collator.compare);apfeläpfelåreölzebraSensitivity check: with base, a vs á is 0; with accent, it is -1.
English and German put these accented words near their base letters in this list. Change to sv to see a different alphabet order.
Pick how closely your phone should compare contact names. For search, a, A, and á can count as equal. For spelling, the accent can matter.
- In real life: Ignore case and accents
- In JavaScript:
sensitivity: "base" - In real life: Keep accent differences
- In JavaScript:
sensitivity: "accent" - In real life: Keep case differences
- In JavaScript:
sensitivity: "case"orvariant
Where the analogy stops: A contact list also sorts names. JavaScript follows the locale and sensitivity you choose.
Intl.Collator is the reusable comparator
SORTERlocaleCompare is convenient for one or two comparisons. For a big sort, create one Intl.Collator and reuse its compare method. That avoids re-explaining the same locale and options for every pair.
| Option | What it changes | Example |
|---|---|---|
sensitivity: "base" | Ignore case and accents for many languages | a, A, and á compare equal in English |
sensitivity: "accent" | Ignore case, but keep accent differences | a and A match; a and á differ |
sensitivity: "case" | Ignore accents, but keep case differences | Useful only when case matters more than accents |
sensitivity: "variant" | Most precise common comparison | Case, accents, and variants can all matter |
numeric: true | Compare digit runs as numbers | item9 sorts before item10 |
caseFirst | Ask whether upper or lower case sorts first | upper, lower, or false |
"Zebra" < "apple"["banana", "apple", "Cherry"].sort()names.sort((a, b) => a.localeCompare(b, "de"))names.sort(new Intl.Collator("sv").compare)new Intl.Collator("en", { numeric: true })"a".charCodeAt(0) > "Z".charCodeAt(0)
Sort each comparison into the rule it uses.
numeric: true is the file-name option people notice first. Without it, item10 often sorts before item9 because 1 comes before 9. With it, digit runs compare as numbers.
Case conversion can be locale-aware too
INTERACTIVEtoUpperCase() and toLowerCase() are useful, but not every language treats letters like English. Locale-aware forms let you name the rules: toLocaleUpperCase(locale) and toLocaleLowerCase(locale).
"i".toLocaleUpperCase("tr"); // "İ""I".toLocaleLowerCase("tr"); // "ı""ß".toLocaleUpperCase("de"); // "SS"İıSSTurkish has dotted and dotless I, and German ß uppercases to SS in JavaScript. Locale-aware case methods let you name the rules you want.
tr and de locales so the result is not guessed from the reader’s browser settings.This matters most when you store or compare folded text. A Turkish user can reasonably expect dotted i and dotless ı behavior. A German user can reasonably expect ß to uppercase to SS in JavaScript. Do not invent your own case table.
Case-insensitive search needs a clear policy
INTERACTIVEThe beginner pattern is haystack.toLowerCase().includes(needle.toLowerCase()). It is fine for simple English labels. Its limits are accents, locale-specific case rules, and Unicode normalization. For equality, a collator with sensitivity: "base" is often clearer.
const collator = new Intl.Collator("en", { sensitivity: "base" }); collator.compare("Résumé", "resume") === 0; // true"Résumé" === "resume"; // falseFor an includes-style search, Intl.Collator does not directly answer “does this folded string contain that folded substring?” A common small-list approach is to normalize to NFD, remove combining marks, then lowercase. Normalize user input too.
function foldForSearch(value) { return value .normalize("NFD") .replace(/[\u0300-\u036f]/g, "") .toLocaleLowerCase("en");} const names = ["José", "Soren", "Søren", "Zoë", "İpek", "Ipek"];const matches = names.filter((name) => foldForSearch(name).includes(foldForSearch(query)),);Folded query: jose
JoséThe query folds to "jose", so 1 name match.
Where you’ll use this
You will use text comparison anywhere users scan a list: contact names, countries, product titles, table columns, command palettes, and file names. The safest pattern is: keep the original string for display, build a comparison or search helper for the task, and name the locale/options explicitly.
const collator = new Intl.Collator(userLocale, { sensitivity: "base", numeric: true,}); const visibleProducts = products .filter((product) => matchesSearch(product.name, query)) .sort((a, b) => collator.compare(a.name, b.name));Reuse one Intl.Collator for a big sort. Creating it once beside the sort is usually clearer and cheaper than calling localeCompare with the same options thousands of times.
Common misconceptions
- “Alphabetical” is universal. Different languages place
ä,å, andödifferently. sort()knows my users. Plain sort uses code-unit string order unless you pass a comparator.- The sign from
localeCompareis always -1 or 1. Only negative, zero, or positive is guaranteed. - Lowercasing both sides solves search forever. It misses accent and locale choices.
- Equal for sorting means identical strings. With
sensitivity: "base", different strings can compare as equal.
Practice exercises
5 EXERCISESRun the code mentally before checking. Type the exact order printed by the default sort.
const words = ["banana", "apple", "Cherry", "éclair"];
console.log(words.sort().join(", "));The default order is Cherry, apple, banana, éclair: capital C first, then lowercase words, then é after these ASCII letters.
Use German collation rules and keep the original array safe by sorting a copy.
const names = ["Özil", "Oden", "Änne", "Anna"];
const collator = new Intl.Collator("de");
console.log([...names].sort(collator.compare).join(", "));const names = ["Özil", "Oden", "Änne", "Anna"];
const collator = new Intl.Collator("de");
console.log([...names].sort(collator.compare).join(", "));German collation places Änne near Anna and Özil near Oden, so the printed order is Anna, Änne, Oden, Özil.
Make the file names sort the way a file picker should: 1, 2, 10.
const files = ["file10", "file2", "file1"];
const collator = new Intl.Collator("en", { numeric: true });
console.log([...files].sort(collator.compare).join(", "));const files = ["file10", "file2", "file1"];
const collator = new Intl.Collator("en", { numeric: true });
console.log([...files].sort(collator.compare).join(", "));The collator reads digit runs as numbers, so file2 comes before file10.
Fill in the fold-and-filter strategy, then predict which displayed name matches zoe.
function fold(value) {
return value.normalize("NFD").replace(/[\u0300-\u036f]/g, "").toLowerCase();
}
const names = ["Jose", "José", "Zoë"];
console.log(names.filter((name) => fold(name).includes(fold("zoe"))).join(", "));function fold(value) {
return value.normalize("NFD").replace(/[\u0300-\u036f]/g, "").toLowerCase();
}
const names = ["Jose", "José", "Zoë"];
console.log(names.filter((name) => fold(name).includes(fold("zoe"))).join(", "));Zoë becomes zoe, so query zoe matches it even though the display keeps the accent.
The bug lowercases both strings but still treats accents as different. Replace it with a collator equality check.
// Bug: accents are treated as different.
console.log("Résumé".toLowerCase() === "resume".toLowerCase());const collator = new Intl.Collator("en", { sensitivity: "base" });
console.log(collator.compare("Résumé", "resume") === 0);The collator says the two strings compare equal at base sensitivity, so compare(...) === 0 prints true.
Check your understanding
7 QUESTIONSQuestion 1 of 7What does JavaScript’s
<use when both values are strings?Choose an answer to see the explanation.
Question 2 of 7What does this print?
Read the code, then predictconsole.log("Zebra" < "apple");Choose an answer to see the explanation.
Question 3 of 7Which tool is best for sorting thousands of displayed names with the same rules?
Choose an answer to see the explanation.
Question 4 of 7What does natural sorting print?
Read the code, then predictconst c = new Intl.Collator("en", { numeric: true }); console.log(["item10", "item9", "item1"].sort(c.compare).join(", "));Choose an answer to see the explanation.
Question 5 of 7What does this print in Turkish?
Read the code, then predictconsole.log("i".toLocaleUpperCase("tr"));Choose an answer to see the explanation.
Question 6 of 7For case- and accent-insensitive equality, which comparison is clearest?
Choose an answer to see the explanation.
Question 7 of 7What does the accent-insensitive fold print?
Read the code, then predictconst folded = "Résumé".normalize("NFD").replace(/[̀-ͯ]/g, "").toLowerCase(); console.log(folded);Choose an answer to see the explanation.
Key takeaways
- Default string comparison is UTF-16 code-unit comparison: fast, stable, and often alien to humans.
- Use
localeComparefor one-off locale-aware comparisons. - Use one
Intl.Collatorwhen sorting many items with the same locale and options. - Locale-aware case conversion matters for languages such as Turkish, and
ßcan expand toSS. - Case-insensitive search needs a policy for normalization, accents, and locale.
Comparing text well means choosing the rule your user expects, not the rule JavaScript can apply fastest.
Up next: Destructuring arrays.