Regex performance & safety
Learn how JavaScript regex backtracking causes slow matches and ReDoS, then use sticky tokenizers, safer patterns, limits, and parsers.
- 01Trace backtrackingUse a tiny matcher to count choices and recognize exponential growth before timing hurts users.
- 02Defend against ReDoSCombine rewrites, input limits, workers, linting, and safer engines when patterns touch untrusted input.
- 03Tokenize and choose parsersUse sticky
yandlastIndexfor lexers, and reach for DOM, URL, JSON, date, and string APIs when they fit better.
Safety before cleverness
Regular expressions are a compact way to describe text, but JavaScript’s normal engine is a backtracking engine. When a pattern leaves too many choices open, one short string can make the engine revisit branch after branch before it can say “no.”
Catastrophic backtracking is a regex match that explores a rapidly growing number of choices before it finishes. ReDoS is the production failure that happens when an attacker uses that slow match to deny service to other work.
This is the last lesson in the Regular expressions module. It builds on patterns and flags, character classes, greedy and lazy quantifiers, groups and backreferences, and anchors and lookaround. The goal is not to fear regex. The goal is to know when a pattern is safe enough for real input.
| Question | Safe direction | Warning sign |
|---|---|---|
| Can the same characters be split many ways? | Make the choices unambiguous. | Nested quantifiers or alternatives that overlap. |
| Can the input be user-controlled? | Cap length before matching. | Patterns run on request bodies, logs, or search boxes. |
| Is the text a real language or format? | Use a parser or platform API. | HTML, URLs, JSON, dates, or full email validation. |
You will write and step through a tiny matcher, compare counted growth for two dangerous shapes, build a sticky tokenizer, and decide when a regex should be replaced by APIs such as URL objects, JSON.parse, the DOM tree, and string methods.
How a backtracking engine explores choices
STEP THROUGHA backtracking engine tries a path, and if a later piece fails, it walks back to the most recent choice and tries a different path. Greedy quantifiers create choices because they first take as much as possible, then give characters back when the suffix cannot match.
Imagine every quantifier choice as a door in a maze. The first route may look promising, but a locked exit makes you walk back and try the next door. If each door leads to more doors, a tiny maze becomes a long search.
- In real life: Take the longest hallway first
- In JavaScript: A greedy
a+consumes everyait can - In real life: A locked exit forces you back
- In JavaScript: The
$anchor fails when the next character is! - In real life: Try a shorter hallway split
- In JavaScript: The engine backtracks and chooses a smaller
stop
Where the analogy stops: A real regex engine is optimized machine code with caches and shortcuts. The maze is only a mental model for why ambiguous choices grow.
The experiment below implements a deliberately tiny matcher for the shape (a+)+$. It only understands a run of a characters followed by the end anchor, and it tests from the start of the string. That small scope is useful: the matcher counts every state and then compares the final true-or-false result with the native JavaScript regex.
Step through a tiny backtracking matcher for a subset of (a+)+$. It counts choices so we can study growth without timing the real engine.
script
function matchNested(input) { let steps = 0; function repeatFrom(index) { steps++; if (index === input.length) return true; if (input[index] !== "a") return false; let end = index; while (input[end] === "a") end++; for (let stop = end; stop > index; stop--) { if (repeatFrom(stop)) return true; } return false; } const matched = repeatFrom(0); return { matched, steps };} const result = matchNested(input);console.log(result.matched);console.log(result.steps);console.log(/^(a+)+$/.test(input));The key line is the loop over stop. For "aaa!", the inner a+ can treat the three letters as aaa, aa then a, a then aa, or three separate chunks. All fail because ! is not the end of the string.
function matchNested(input) { let steps = 0; function repeatFrom(index) { steps++; if (index === input.length) return true; if (input[index] !== "a") return false; let end = index; while (input[end] === "a") end++; for (let stop = end; stop > index; stop--) { if (repeatFrom(stop)) return true; } return false; } const matched = repeatFrom(0); return { matched, steps };} const input = "aaa!";const result = matchNested(input);console.log(result.matched);console.log(result.steps);console.log(/^(a+)+$/.test(input));(a+)+$32engine result: false(a|aa)+$20engine result: false| length | (a+)+$ | (a|aa)+$ |
|---|---|---|
| 1 | 2 | 2 |
| 2 | 4 | 4 |
| 3 | 8 | 7 |
| 4 | 16 | 12 |
| 5 | 32 | 20 |
| 6 | 64 | 33 |
| 7 | 128 | 54 |
The numbers are from the lesson's tiny matcher, not from timing the browser's regex engine. The native engine is used only for the bounded true-or-false agreement check.
The page never benchmarks a catastrophic native regex. The slider counts choices in a bounded teaching matcher and keeps inputs tiny. Tests prove the matcher’s accept or reject result agrees with JavaScript on those bounded inputs.
The three shapes that create too many paths
SORT ITMost ReDoS reviews start by naming the shape. A pattern can be long and safe, or short and dangerous. The risky part is not “regex” by itself; it is ambiguity that grows with input.
const examples = [ "nested quantifier: ^(a+)+$", "overlap: ^(a|aa)+$", "ambiguous neighbors: ^a*a*$",];console.log(examples.length);console.log(examples[0].includes("nested"));If you can name the shape, the rewrite becomes much easier: remove nesting, make alternatives disjoint, or add a real boundary.
^(a+)+$^(.*)+$^(a|aa)+$^(cat|catalog)+$^a*a*$^\w+\w+$/<div>.*<\/div>/for real HTML/{.*}/for JSON
Sort each pattern by the shape that makes it risky. Some cards are not regex jobs at all.
| Root cause | Risky shape | Safer direction |
|---|---|---|
| Nested quantifier | ^(a+)+$ | ^[a]+$ or a loop that checks every character once |
| Overlapping alternation | ^(a|aa)+$ | Make alternatives disjoint, or parse one token at a time |
| Ambiguous neighbors | ^a*a*$ | Use one quantifier, add a real separator, or use a character class boundary |
| Untrusted length | Any complex pattern | Reject or slice oversized input before matching |
JavaScript does not have possessive quantifiers such as a++ or atomic groups such as (?>...). Some regex flavors use those to prevent backtracking into a piece, but JavaScript cannot. Prefer unambiguous character classes, disjoint alternatives, whole-string anchors when you mean the whole string, and input length limits.
function acceptsSlug(slug) { if (slug.length > 32) return false; return /^[a-z]+(?:-[a-z]+)*$/.test(slug);} console.log(acceptsSlug("regex-performance"));console.log(acceptsSlug("regex--performance"));console.log(acceptsSlug("a".repeat(40)));truefalsefalseThis example is intentionally boring: one clear separator, whole-string anchors, and a length cap before matching. Boring is a security feature.
ReDoS: when one crafted string blocks everyone
PRODUCTIONJavaScript runs user code on a single main thread in the browser and on an event loop in Node. A slow regex is CPU work. While that match is running, it does not yield to clicks, timers, rendering, or other server requests. Review the event loop lesson if you want the scheduling model behind that statement.
On the server, a single request body, header, username, or log line can trigger a slow regex. If the process is busy matching, every other request handled by that event loop waits. That is why regex performance is a security and reliability topic.
- Stack Overflow, July 2016. A malformed post with roughly 20,000 whitespace characters triggered heavy backtracking in a whitespace-trimming regex and the site was down for 34 minutes. The postmortem described the simplified case as
\s+$and noted it was quadratic rather than classic exponential, but it was still enough regex work to take down production. Source: ADTmag’s summary of the StackStatus postmortem. - Cloudflare, July 2019. A WAF rule deployment included a regex that drove CPU to 100% across HTTP/HTTPS serving cores. Cloudflare reported a 27-minute global outage, 13:42 to 14:09 UTC, with traffic dropping by 82% at peak. Source: Cloudflare’s postmortem.
| Defense | Where it helps | How to apply it |
|---|---|---|
| Length caps | Fast first line of defense | Reject or truncate before the regex runs; document the limit. |
| Rewrite risky shapes | Most durable fix | Remove nested quantifiers, overlapping alternatives, and ambiguous adjacent quantifiers. |
| Worker timeout | Browser isolation | Run untrusted patterns in a Web Worker and call terminate() if a deadline passes. |
| Linters and scanners | Build-time warnings | Use eslint-plugin-regexp and tools such as safe-regex with their false-positive limits in mind. |
| RE2 | Runtime guarantee | Node bindings use a linear-time engine, but unsupported patterns include backreferences and lookaround. |
In the browser, isolate untrusted patterns in a Web Worker so the page can keep responding. A timeout cannot interrupt a regex already running on the main thread, but it can terminate a separate worker.
const worker = new Worker("regex-worker.js");const timeout = setTimeout(() => { worker.terminate(); console.log("regex worker stopped");}, 50); worker.onmessage = (event) => { clearTimeout(timeout); console.log(event.data.matched);}; worker.postMessage({ pattern: "^(a+)+$", input: "aaaa!" });Tooling helps before runtime. eslint-plugin-regexp finds regex mistakes, style issues, and optimization hints. safe-regex flags many potentially exponential patterns but documents false positives and false negatives. For Node services, RE2 bindings use a linear-time engine for supported patterns, but backreferences and lookaround are not supported. V8 also documents an experimental non-backtracking engine and fallback flags; do not assume that protection is on by default in ordinary JavaScript.
The sticky flag and lastIndex
STEP THROUGHThe g flag and the sticky y flag both use lastIndex. The difference is where the next match may start. A global regex may scan forward from lastIndex. A sticky regex must match exactly at lastIndex or fail.
| Behavior | g global | y sticky |
|---|---|---|
| Where matching may start | At lastIndex, but it may scan forward to a later match | Exactly at lastIndex; no scanning ahead |
| Success | Sets lastIndex after the match | Sets lastIndex after the match |
| Failure | Resets lastIndex to 0 | Resets lastIndex to 0 |
| Best use | Find every occurrence in a string | Build lexers and tokenizers that must not skip characters |
That exact cursor is perfect for a small lexer. The tokenizer below consumes optional whitespace and then one identifier, number, or equals sign. If the next character is not a token, it fails at that cursor instead of hiding the problem by scanning ahead.
Step through a sticky y tokenizer. Watch lastIndex become the cursor that every token must start from.
script
const input = "let count = 3";const parts = []; while (token.lastIndex < input.length) { const before = token.lastIndex; const match = token.exec(input); if (!match) throw new Error("stuck at " + before); parts.push(match[1]);} console.log(parts.join(" | "));console.log(token.lastIndex);const text = "xxcat";const global = /cat/g;global.lastIndex = 1;const foundGlobal = global.exec(text);console.log(foundGlobal.index);console.log(global.lastIndex); const sticky = /cat/y;sticky.lastIndex = 1;console.log(sticky.exec(text));console.log(sticky.lastIndex);25null0The global regex starts at index 1 and finds cat at index 2. The sticky regex starts at index 1, refuses to skip x, returns null, and resets lastIndex to 0.
lastIndex; sticky makes that index exact.A failed exec or test with g or y resets lastIndex to 0. Save the old cursor before matching if you need a helpful “unexpected character at index …” error.
When not to use a regex
CHOOSE APIA regex is great for regular text shapes: slugs, simple tokens, a known prefix, or a small capture in a line. It is the wrong first tool when the platform already has a parser that understands the format.
| Task | Regex risk | Use instead |
|---|---|---|
| HTML | Fragile: nesting and malformed markup break text patterns | DOMParser or DOM methods |
| URLs | Easy to miss escaping, base URLs, and query decoding | new URL() and URLSearchParams |
| JSON | Regex cannot reliably handle nested strings and escapes | JSON.parse |
| Full validation is a protocol and delivery problem | Use a simple shape check, then confirm by sending | |
| Dates | Locales, calendars, and time zones are not text trivia | Temporal when available, plus Intl for formatting |
| Simple text | Often harder to read than a method call | includes, startsWith, endsWith, split |
const doc = new DOMParser().parseFromString( "<article><h1>Regex</h1><p>Use a parser.</p></article>", "text/html",);console.log(doc.querySelector("h1")?.textContent);console.log(doc.querySelector("p")?.textContent);const page = new URL("/search?q=regex", "https://example.com");console.log(page.searchParams.get("q"));console.log(JSON.parse('{"safe":true}').safe);console.log("lesson.regex.md".endsWith(".md"));console.log(new Intl.DateTimeFormat("en-US", { month: "short", timeZone: "UTC",}).format(new Date("2026-09-26T00:00:00Z")));Email deserves special caution. A giant “perfect email regex” is rarely user-friendly or complete. For most products, keep the client-side check simple, then confirm ownership by sending an email. Dates are similar: use Temporal when available for date math and Intl for display instead of trying to encode calendars and time zones in a pattern.
Common misconceptions
- “Lazy quantifiers are safe.” Lazy quantifiers still backtrack. They just try the shortest choice first.
- “Anchors fix performance.” Anchors reduce where a match can start, but they do not remove ambiguous choices inside the pattern.
- “Short test strings prove a regex is safe.” Dangerous patterns often look fine until a crafted failure string grows by a few characters.
- “Static tools prove safety.” Linters and scanners are useful warnings, but they can miss patterns or flag safe ones. Treat them as one layer.
- “Regex is always faster than string methods.” For simple contains, prefix, suffix, or split work, string methods are clearer and often faster.
| Belief | Better rule | Why |
|---|---|---|
.* is harmless when the input is small | Ask who controls the input and cap length | Attackers choose edge cases, not demo cases. |
| A clever regex is a parser | Use the parser for real formats | HTML, JSON, URLs, and dates already have syntax rules. |
| A worker makes the regex safe | A worker only isolates the damage | You still need a timeout and a pattern review. |
Practice exercises
5 EXERCISESRead the bounded matcher and type the number printed.
function matchNested(input) {
let steps = 0;
function repeatFrom(index) {
steps++;
if (index === input.length) return true;
if (input[index] !== "a") return false;
let end = index;
while (input[end] === "a") end++;
for (let stop = end; stop > index; stop--) {
if (repeatFrom(stop)) return true;
}
return false;
}
const matched = repeatFrom(0);
return { matched, steps };
}
const result = matchNested("aaa!");
console.log(result.steps);The tiny matcher visits eight states before every split of aaa fails against the final !.
Predict the exact token string printed by the sticky lexer.
const token = /\s*([A-Za-z_]\w*|\d+|=)/y;
const input = "let count = 3";
const parts = [];
while (token.lastIndex < input.length) {
const match = token.exec(input);
parts.push(match[1]);
}
console.log(parts.join(" | "));Sticky matching walks left to right and captures let, count, =, and 3, so the joined output is let | count | = | 3.
After the failed match, what number is printed for lastIndex?
const sticky = /cat/y;
sticky.lastIndex = 1;
console.log(sticky.exec("xxcat"));
console.log(sticky.lastIndex);/cat/y cannot match at index 1, so exec returns null and resets lastIndex to 0.
Type the two booleans printed by the safe slug checker.
function acceptsSlug(slug) {
if (slug.length > 32) return false;
return /^[a-z]+(?:-[a-z]+)*$/.test(slug);
}
console.log(acceptsSlug("regex-performance"));
console.log(acceptsSlug("regex--performance"));regex-performance passes. regex--performance fails because the non-capturing group requires a hyphen followed by letters, not another hyphen.
Which built-in object should you use to parse URLs instead of a regex?
const page = new URL("/search?q=regex", "https://example.com");
console.log(page.searchParams.get("q"));const page = new URL('/search?q=regex', 'https://example.com');Use the URL object and searchParams for URL parsing instead of a hand-written regex.
Check your understanding
7 QUESTIONSQuestion 1 of 7What is catastrophic backtracking?
Choose an answer to see the explanation.
Question 2 of 7What does the tiny matcher print for
aaa!?Read the code, then predictfunction matchNested(input) { let steps = 0; function repeatFrom(index) { steps++; if (index === input.length) return true; if (input[index] !== "a") return false; let end = index; while (input[end] === "a") end++; for (let stop = end; stop > index; stop--) { if (repeatFrom(stop)) return true; } return false; } const matched = repeatFrom(0); return { matched, steps }; } const result = matchNested("aaa!"); console.log(result.steps);Choose an answer to see the explanation.
Question 3 of 7Which risky shape does
^(a|aa)+$contain?Choose an answer to see the explanation.
Question 4 of 7What does this sticky miss leave in
lastIndex?Read the code, then predictconst sticky = /cat/y; sticky.lastIndex = 1; console.log(sticky.exec("xxcat")); console.log(sticky.lastIndex);Choose an answer to see the explanation.
Question 5 of 7Why is ReDoS especially dangerous in Node.js servers?
Choose an answer to see the explanation.
Question 6 of 7Which statement about RE2 is accurate?
Choose an answer to see the explanation.
Question 7 of 7What should parse this JSON string?
Read the code, then predictconst text = '{"ok":true}'; console.log(JSON.parse(text).ok);Choose an answer to see the explanation.
Key takeaways
- Backtracking engines can revisit many equivalent choices before a match fails.
- Nested quantifiers, overlapping alternation, and ambiguous neighboring quantifiers are the main danger signs.
- ReDoS blocks real work because regex matching is CPU work on the JavaScript thread or Node event loop.
- Sticky
ymakeslastIndexan exact lexer cursor; failed global or sticky matches reset it to 0. - Use parsers and platform APIs for HTML, URLs, JSON, dates, email confirmation, and simple string checks.
Remember the one-liner.
Safe regex work is about reducing choices, bounding input, and choosing a parser when text has real structure.
That wraps the Regular expressions module. Next comes typed arrays in Binary data & memory, where the focus shifts from text patterns to bytes, buffers, and memory views.