Character classes & sets
Learn how JavaScript regex character classes match ASCII digits, words, spaces, Unicode properties, v-flag sets, and safe user input.
- 01Read shorthand classes preciselyExplain what
\d,\w,\s, their uppercase opposites, and dot really match. - 02Choose sets or Unicode propertiesUse ranges, negated sets,
\p{L},\p{N}, scripts, emoji properties, and\P{...}intentionally. - 03Build safer dynamic patternsEscape metacharacters correctly with a helper or ES2025
RegExp.escape, and know when thevflag is stricter.
What character classes mean
A regular expression is a sequence of positions to match. A character class describes what one position may contain: one digit, one word character, one letter from any script, one character from a bracket set, or one character outside a set.
A character class is a regex atom that matches one character from a set of allowed characters. Shorthands such as \d, bracket sets such as [a-z], Unicode properties such as \p{L}, and v-flag set operations are all ways to describe that set.
The previous lesson, Patterns & flags, covers literals, flags, test, match, and replace. This lesson stays in one lane: what counts as a character. For text background, review strings and Unicode & strings.
| Syntax | What it means | Best use |
|---|---|---|
\w | ASCII word characters: A-Z, a-z, 0-9, and _. | Use for slugs, simple identifiers, and machine formats. |
\p{L} | Unicode letters across scripts, when used with u or v. | Use for international names and text. |
u flag | Unicode-aware code point matching and property escapes. | Needed for \p{...} and sane astral-character matching. |
v flag | Unicode sets: &&, --, nested classes, and string properties. | Use when you need set algebra or \p{RGI_Emoji}. |
Built-in classes: \d, \w, \s, and dot
STEP THROUGHShorthand classes are compact, but their names are easy to overread. In JavaScript, \d is ASCII digits only, \w is the ASCII word set, and \s is broader because it includes Unicode whitespace. Uppercase versions negate the class: \D, \W, and \S.
Think of a character class as a bouncer at one door. The bouncer does not judge the whole sentence, just the next character trying to enter this one position.
- In real life: The bouncer reads one printed guest list
- In JavaScript: The regex engine checks one character class
- In real life: Someone either is or is not on that list
- In JavaScript: A character either matches
\d,\w, or the chosen set - In real life: A different door can have a different list
- In JavaScript: Switching from
\wto\p{L}changes the allowed set
Where the analogy stops: A bouncer checks people, not code points. Regexes also have flags, escaping rules, and string properties that can go beyond this one-door analogy.
The lab records the checks. Notice that \d skips ٣, while \p{N} accepts it; \w accepts underscore, but not é; and dot needs the s flag for newlines plus u or v for astral characters to count as one match.
Choose a built-in class, predict which labels survive, then step through the real checks.
script
["ASCII 9", "9"], ["Arabic ٣", "٣"], ["A", "A"], ["underscore", "_"], ["space", " "],]; const labels = (regex) => samples .filter(([, value]) => regex.test(value)) .map(([label]) => label) .join(", "); console.log(labels(/\d/u));console.log(labels(/\D/u));console.log(/\p{N}/u.test("٣"));const sample = "A é β 9 ٣ _ -\u00A0.\n☕ 😀";const matches = sample.match(/\d/gu) ?? [];console.log(matches.join(" | "));Match list: 9
AskipsAspaceskipsspaceéskipséspaceskipsspaceβskipsβspaceskipsspace9matches9spaceskipsspace٣skips٣spaceskipsspace_skips_spaceskipsspace-skips-NBSPskipsNBSP.skips.\nskips\n☕skips☕spaceskipsspace😀skips😀ASCII digits only: 0 through 9. The match list comes from the displayed global regular expression.
\wTreat \w as the ASCII word class. Unicode-aware, case-insensitive matching can case-fold a few compatibility characters, but \w is still not a Unicode letter class. Use \p{L} when the rule is “letters from human writing systems.”
Sets and ranges: [abc], [a-z0-9], and [^...]
SORT ITBrackets make your own one-character set. [abc] means one a, b, or c. Ranges use a hyphen between endpoints: [a-z0-9] means one lowercase ASCII letter or ASCII digit. A caret immediately after [ negates the set: [^a-z0-9] means one character outside that set.
console.log("cab".match(/[abc]/g).join(""));console.log("B7!".match(/[A-Z0-9]/g).join(""));console.log("B7!".match(/[^A-Z0-9]/g).join(""));console.log("a-z".match(/[a\-z]/g).join(""));Escaping rules change inside brackets. Outside a set, ., +, ?, parentheses, braces, |, anchors, and backslash can be syntax. Inside a classic set, the big ones are ], backslash, ^ at the start, and - between endpoints. Put a literal hyphen first or last, or escape it as \-.
\d[a-z0-9]\p{L}\p{Script=Greek}[\p{L}--[aeiou]]\p{RGI_Emoji}
Sort each pattern by the syntax family it belongs to. This is about the class itself, not about quantifiers or groups.
Unicode properties with \p{...} and \P{...}
STEP THROUGHUnicode property escapes require the u or v flag. Lowercase \p means “has this property.” Uppercase \P means “does not have this property.” Useful properties include \p{L} for letters, \p{Lu} for uppercase letters, \p{N} for numbers, \p{Script=Greek} for Greek script, and \p{Emoji_Presentation} for code points that default to emoji presentation.
Unicode property escapes require u or v. Pick a property family and step through exactly which text survives.
script
console.log(text.match(/\p{L}/gu).join(" "));console.log(text.match(/\p{N}/gu).join(" "));console.log(text.match(/\P{L}/gu).filter((ch) => ch !== " ").join("|"));Properties are not a replacement for product rules. For international names, a pattern may need letters, combining marks, spaces, apostrophes, and hyphens. For usernames, you may intentionally want ASCII only. The regex should say the policy, not guess it.
The v flag: set operations and string properties
INTERACTIVEThe v flag, also exposed as unicodeSets, builds on Unicode-aware regexes. It adds set difference with --, intersection with &&, nested classes, and string properties. A string property can match more than one code point, such as an RGI emoji sequence, as a single match.
const text = "alphabet β123";const consonantsAndGreek = text.match(/[\p{L}--[aeiou]]/gv) ?? [];console.log(consonantsAndGreek.join(""));lphbtβThe v flag reads [\p{L}--[aeiou]] as letters minus lowercase ASCII vowels. Greek beta remains because it is a letter and not one of the subtracted vowels.
v flag. The string-property option uses emoji because string properties are mainly useful for emoji sequences today.Because v gives punctuation new meaning inside sets, it is stricter about escaping. For example, a literal | inside a v-mode set must be escaped. The next snippet throws on purpose so you can read the error and fix it to [a\|b].
v-mode escaping error// In v mode, | is reserved inside a character class.new RegExp("[a|b]", "v");Escaping user input and RegExp.escape
SAFE DYNAMIC REGEXWhen a learner types a search term, that text is data. If you place it directly inside new RegExp(term), characters such as +, [, and ? become regex syntax. A safe highlighter escapes the term before building the regex.
function escapeForRegExp(input) { return input.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");} const userText = "a+b [test]";const escaped = escapeForRegExp(userText);const regex = new RegExp(escaped, "g"); console.log(escaped);console.log("find a+b [test] here".replace(regex, "<mark>$&</mark>"));ES2025 adds RegExp.escape, which is more complete than the classic helper. MDN browser-compat data lists support in Chrome 136+, Firefox 134+, Safari 18.2+, and Node.js 24+. The site tests run on Node 22, which does not have it, so server-side Node 22 code still needs the helper or a polyfill.
const userText = "a+b [test]";const escaped = RegExp.escape(userText);const regex = new RegExp(escaped, "g"); console.log(escaped);console.log(regex.test("find a+b [test] here"));Practical uses
Character classes show up in validation, cleanup, search, and tokenizing. Start with the rule in plain language: “ASCII username,” “digits from a phone field,” “letters in a human name,” or “literal search term.” Then choose the class that says exactly that.
| Problem | Pattern direction | Why |
|---|---|---|
| Username policy | /^[A-Za-z0-9_-]+$/ or a Unicode property version | Decide whether the policy is ASCII-only before reaching for \w. |
| Phone digits | value.replace(/\D/g, "") | Good for ASCII phone input; use \p{N} only if you truly accept digits from many scripts. |
| International names | /^[\p{L}\p{M}' -]+$/u | Letters and marks are broader than \w; validation still needs product rules. |
| Search-term highlight | new RegExp(RegExp.escape(term), "gi") | User text is data, not regex syntax. |
- Usernames:
^[A-Za-z0-9_-]+$is honest when the policy is ASCII. - Phone digits:
value.replace(/\D/g, "")strips non-ASCII digits from common phone input. - Stripping punctuation:
text.replace(/[^\p{L}\p{N}\s]/gu, "")keeps letters, numbers, and spaces. - International names: combine
\p{L}, marks, spaces, and punctuation only when your product accepts them. - Search highlighting: escape the user term, then build the dynamic regex.
Common misconceptions
- “
\dmeans any digit.” In JavaScript it means ASCII0through9. Use\p{N}for Unicode numbers. - “
\wmeans any word letter.” It is the ASCII word set, not an international name validator. - “Dot means absolutely everything.” Dot skips line terminators unless the
sflag is present, and astral characters needuorvto count as one code point. - “A bracket set matches a word.” A set matches one character. Quantifiers in the next lesson decide how many repetitions are allowed.
- “The
vflag is justuwith a new name.” It adds set operations, string properties, and stricter parsing. - “User input is safe because it came from a text box.” Text can contain regex syntax. Escape it before building a pattern.
| Choice | Use it when | Do not use it when |
|---|---|---|
\w | The policy is ASCII identifiers or simple slugs. | You mean letters from many scripts. |
\p{L} | You mean Unicode letters and have u or v. | You need a tight ASCII machine format. |
u | You need Unicode-aware code points and properties. | You need &&, --, nested classes, or string properties. |
v | You need set algebra or \p{RGI_Emoji}. | A simpler u property or bracket set is clearer. |
Practice exercises
5 EXERCISESType the exact first line printed by this snippet.
const values = ["9", "٣"];
console.log(values.filter((value) => /\d/u.test(value)).join("|"));
console.log(/\p{N}/u.test("٣"));Only 9 passes \d. The second log would print true because \p{N} accepts the Arabic-Indic digit.
The ASCII-only username regex rejects useful names. Which core class should replace \w when the rule is “letters from many scripts”?
const names = ["maria", "maria-é", "δοκιμή"];
const asciiOnly = /^[\w-]+$/u;
console.log(names.filter((name) => asciiOnly.test(name)).join(","));const international = /^[\p{L}\p{N}_-]+$/u;\p{L} accepts letters such as é and Greek letters. Keep _ and - explicit because they are product-rule punctuation, not letters.
Type the exact string printed by the corrected filter.
const names = ["maria", "maria-é", "δοκιμή"];
const international = /^[\p{L}\p{N}_-]+$/u;
console.log(names.filter((name) => international.test(name)).join(","));All three names pass the Unicode property version, so the joined output is maria,maria-é,δοκιμή.
v-flag set differencePredict the output from the `v`-flag difference.
const text = "a b c x y z";
console.log(text.match(/[[a-z]--[aeiou]]/gv).join(""));The pattern keeps lowercase ASCII consonants, so a b c x y z becomes bcxyz.
A user searches for a+b. Which ES2025 method should you call before building the regex?
const term = "a+b";
const unsafe = new RegExp(term);
console.log(unsafe.test("aaab"));const regex = new RegExp(RegExp.escape(term), "g");RegExp.escape turns the search term into regex-safe literal source before new RegExp parses it.
Check your understanding
8 QUESTIONSQuestion 1 of 8What does
\dmatch in JavaScript?Choose an answer to see the explanation.
Question 2 of 8What does the word-class snippet print?
Read the code, then predictconst values = ["é", "_", "A"]; console.log(values.filter((value) => /\w/u.test(value)).join(","));Choose an answer to see the explanation.
Question 3 of 8What does dot print for this astral-character example?
Read the code, then predictconsole.log("😄".match(/./g).length, "😄".match(/./gu).length);Choose an answer to see the explanation.
Question 4 of 8What does the range snippet print?
Read the code, then predictconsole.log("aB7-".match(/[a-z0-9]/g).join(""));Choose an answer to see the explanation.
Question 5 of 8Which output comes from this Greek script check?
Read the code, then predictconsole.log("A β Ж".match(/\p{Script=Greek}/gu).join(""));Choose an answer to see the explanation.
Question 6 of 8What does this
v-flag set difference print?Read the code, then predictconst text = "a b c x y z"; console.log(text.match(/[[a-z]--[aeiou]]/gv).join(""));Choose an answer to see the explanation.
Question 7 of 8Why should a search highlighter escape user input before
new RegExp?Choose an answer to see the explanation.
Question 8 of 8Which runtime fact is true for
RegExp.escape?Choose an answer to see the explanation.
Key takeaways
\dand\ware ASCII shorthands;\sincludes Unicode whitespace.- Dot skips line terminators without
s, and astral characters needuorvto match as one code point. - Bracket sets match one character; ranges and negation describe which one.
- Unicode properties such as
\p{L},\p{Lu},\p{N}, and\p{Script=Greek}requireuorv. - The
vflag adds set operations, nested classes, string properties, and stricter escaping. - Escape user input with a helper or
RegExp.escapebefore building a dynamic regex.
Remember the one-liner.
Say the character set you mean in plain language, then choose the regex syntax that matches exactly that set.
Up next: Quantifiers, where the class you just chose can repeat with +, *, ?, and {n,m}.