Quantifiers: greedy & lazy
Learn JavaScript regex quantifiers, greedy and lazy matching, empty-match traps, and practical patterns for quoted strings, PINs, and cleanup.
- 01Choose the right repeat markerUse
+,*,?,{n},{n,}, and{n,m}without guessing what can be empty. - 02Trace greedy and lazy matchesExplain how a greedy match backtracks and how a lazy match expands just enough.
- 03Avoid quantifier trapsHandle empty matches, quoted text, simple validation lengths, and quantified groups deliberately.
Control how much a pattern matches
A regular expression finds text by moving token by token. A quantifier says how many times the token immediately before it may repeat. Without a quantifier, a token usually means “exactly one.” With a quantifier, it can mean “one or more,” “zero or more,” “optional,” or “between these counts.”
A quantifier is a repeat marker such as +, *, ?, {4}, {2,}, or {2,5}. It applies to the single preceding regex token. Greedy quantifiers try the largest match first; lazy quantifiers add a trailing ? and try the smallest match first.
This lesson stays in the quantifier lane. For creating regexes, flags, and methods, review Patterns & flags. For \d, \w, bracket ranges, and Unicode properties, see Character classes & sets. We will briefly quantify a simple non-capturing group like (?:ab)+, and the next lesson on Groups & backreferences goes deeper.
Imagine dragging a highlighter across repeated words. A note on the highlighter says “mark one or more,” “mark zero or more,” or “mark between two and four.” Quantifiers are those notes. They do not choose the token; they only control how many copies of the previous token can be part of the match.
- In real life: Highlight exactly one word
- In JavaScript: A plain token such as
a - In real life: Keep highlighting while the words match
- In JavaScript:
+or*repeating a token - In real life: Highlight the shortest useful phrase
- In JavaScript: A lazy quantifier such as
+? - In real life: Stop at a boundary color
- In JavaScript: A negated class such as
[^"]*
Where the analogy stops: A human highlighter understands meaning. A regex engine only follows tokens and may backtrack; it does not know that text is HTML, JSON, or a sentence.
+, *, ?, and {n,m}
REAL OUTPUTStart with the three compact quantifiers. a+ means one or more a characters. a* means zero or more a characters. a? means zero or one a. The phrase “zero” is important: both * and ? can let a pattern succeed without consuming that token.
Quantifiers attach to one token. In colou?r, only the u is optional, so it matches color and colour, not colouur. In (?:ab)+, the single preceding token is a non-capturing group, so the whole ab pair repeats.
const words = "color colour colouur";console.log(words.match(/colou?r/g).join(", "));console.log("aaaab".match(/a+b/)?.[0]);console.log("bbb".match(/a*b/)?.[0]);console.log("ababab".match(/(?:ab)+/)?.[0]);Read the output line by line: the optional u? finds two spellings, a+ consumes the whole run of a before b, a* is allowed to match zero a characters before b, and (?:ab)+ repeats a grouped pair. Capturing and non-capturing groups are the next lesson; here the important point is that a group can become one token.
Braces give exact and ranged counts: {4} means exactly four, {2,} means at least two, and {2,4} means from two to four. JavaScript does not support {,4} as “up to four.” In legacy non-Unicode mode new RegExp("a{,3}") behaves like a literal brace pattern, and with u or v it throws a syntax error. Use {0,4} when zero is the lower bound.
{,m}const pins = ["123", "1234", "12345", "12a4"];console.log(pins.filter((pin) => /^\d{4}$/.test(pin)).join(", "));console.log("haaa!".match(/ha{2,}/)?.[0]);console.log("aaaaa".match(/a{2,4}/)?.[0]);console.log(new RegExp("a{,3}").test("a{,3}"));try { new RegExp("a{,3}", "u");} catch (error) { console.log(error.name);}/a+//a*//colou?r//\d{4}//a{2,4}//.+?/
Place each pattern by the minimum it allows. Lazy syntax changes search order, not the minimum count.
Greedy matching and backtracking
STEP THROUGHJavaScript quantifiers are greedy by default. “Greedy” does not mean “wrong.” It means the quantified token first takes as much text as it can, then the engine tries the next token. If the next token cannot match, the engine backtracks: it gives characters back one at a time until the rest of the pattern can finish, or until there is no legal match left.
The classic surprise is /".+"/. The opening quote matches the first quote. Then .+ grabs everything through the end of the string. Finally the closing quote token must match, so the engine backs up from the period to the last quote. The result is one long match, not the first quoted word.
Step through ".+": it grabs to the end first, then backs off only as much as needed for the final quote to match.
script
const engineMatch = text.match(/".+"/)[0];const firstQuote = text.indexOf('"');let cursor = firstQuote + 1;while (cursor < text.length) cursor += 1;while (text[cursor - 1] !== '"') cursor -= 1;const simulated = text.slice(firstQuote, cursor);console.log(engineMatch);console.log(simulated === engineMatch);The recording uses a tiny position simulation for this exact pattern and prints true when the simulated final match agrees with the JavaScript engine. It is a teaching model, not a debugger for every regex feature, but it captures the useful mental model: greedy first, then backtrack only when later tokens demand it.
const text = 'say "one" and "two".';const engineMatch = text.match(/".+"/)[0];const firstQuote = text.indexOf('"');let cursor = firstQuote + 1;while (cursor < text.length) cursor += 1;while (text[cursor - 1] !== '"') cursor -= 1;const simulated = text.slice(firstQuote, cursor);console.log(engineMatch);console.log(simulated === engineMatch);A little backtracking is normal. Nested ambiguous quantifiers can create catastrophic backtracking and ReDoS risks; the Regex performance lesson covers that in depth. Here, just learn to notice when a quantifier has too many ways to match the same text.
Lazy quantifiers and quoted text
INTERACTIVEAdd ? after a quantifier to make it lazy: +?, *?, ??, and {2,5}?. Lazy does not mean “ignore the minimum.” +? still needs one character. It simply starts with the smallest allowed match and expands only if the next token cannot match yet.
Lazy quantifiers begin with the smallest allowed match and expand only when the following token cannot match yet.
script
const lazy = text.match(/".+?"/)[0];const negated = text.match(/"[^"]*"/)[0];let end = text.indexOf('"') + 2;while (text[end] !== '"') end += 1;const simulated = text.slice(text.indexOf('"'), end + 1);console.log(lazy);console.log(negated);console.log(simulated === lazy);Lazy patterns are useful, but they are not always the clearest boundary. For simple quoted strings without escape sequences, /"[^"]*"/ often reads better than /".*?"/: quote, then zero or more non-quote characters, then quote. The boundary is part of the token, not a backtracking strategy.
const text = 'say "one" and "two".';const pattern = /".+"/g;const matches = [...text.matchAll(pattern)].map((match) => match[0]);console.log(matches.join(" | "));- index 4
"one" and "two"
Greedy .+ takes one large match when the text contains more than one quoted part.
| Approach | How it searches | When it fits |
|---|---|---|
| Greedy | Starts with the largest possible span, then backtracks if later tokens fail. | /".+"/ on two quotes returns one long match. |
| Lazy | Starts with the smallest allowed span, then expands only as needed. | /".+?"/ returns the first quoted span. |
| Negated class | Defines the boundary directly by saying what cannot appear inside. | /"[^"]*"/ reads as quote, non-quotes, quote. |
Practical places quantifiers show up
APPLY ITQuantifiers are small, but they show up everywhere: a four-digit PIN, a postcode segment, repeated whitespace cleanup, quoted attributes, and simple Markdown markers. Combine them with character classes from the previous lesson and anchors from a later lesson when you need whole-string validation.
const pins = ["4829", "48290", "48a9"];console.log(pins.filter((pin) => /^\d{4}$/.test(pin)).join(", "));console.log("too much \t space".replace(/\s+/g, " "));const attributes = 'name="Ada" title="Engineer"';console.log(attributes.match(/"[^"]*"/g).join(", "));console.log("Use **bold** here".replace(/\*\*([^*]+)\*\*/g, "<strong>$1</strong>"));Line 2 uses ^\d{4}$ to keep only a whole string of four digits. The ^ and $ anchors are what make it a whole-string check; the Anchors & lookaround lesson covers them properly. Line 3 collapses one or more whitespace characters into one space. Line 5 finds quoted text by repeating non-quote characters. The Markdown line uses a capture group for replacement; the next lesson explains that group in depth.
Common traps
- “
.*means up to the next thing I see.” No. Greedy.*first goes as far as it can, then backtracks only enough for the rest of the pattern. - “
*means there must be a match in the text.” It can match nothing at all, which matters with globalreplaceandmatch. - “
\d{4}validates a four-digit string.” It only finds four digits somewhere. Use anchors for a whole string. - “Quantifying a capture stores every repetition.” JavaScript keeps the last capture for that group. Use a different matching strategy when every repeated part matters.
| Trap | Why it happens | Safer habit |
|---|---|---|
.* in markup-like text | It can cross from the first opening tag to the last closing tag. | Prefer a parser for real HTML; for tiny demos, use a tighter class or a lazy pattern. |
* with a global replacement | It can match the empty string between every character. | Ask whether zero characters should really count as a match. |
| Length validation without anchors | /\d{4}/ finds four digits inside 12345. | Use anchors for whole-string checks; the anchors lesson goes deeper. |
| Quantified capturing group | Only the last capture from the repeated group is kept. | Use a non-capturing group for grouping, or collect all matches another way. |
.* in HTML-ish textconst html = "<strong>one</strong><strong>two</strong>";console.log(html.match(/<strong>.*<\/strong>/)[0]);console.log(html.match(/<strong>.*?<\/strong>/)[0]);console.log(html.match(/<strong>[^<]*<\/strong>/)[0]);This is intentionally HTML-ish text, not a recommendation to parse HTML with regex. The example isolates one trap: the greedy version jumps from the first <strong> to the last </strong>. Real markup needs a parser.
* can match the empty stringconsole.log("abc".replace(/x*/g, "-"));console.log("xx".replace(/x*/g, "-"));console.log("abc".match(/x*/g).length);/x*/g can match zero x characters before a, between letters, and after c. JavaScript advances after an empty global match so it does not loop forever, but your replacement still appears in surprising places.
const capture = "abc".match(/([abc])+/);console.log(capture[0]);console.log(capture[1]);console.log("ababab".match(/(?:ab)+/)[0]);The full match is abc, but the capture is c because that group ran three times and the last run won. When you only need grouping, prefer (?:ab)+; when you need all repeated pieces, use matchAll, a different pattern, or a parser depending on the job.
Practice exercises
5 EXERCISESuRead the three tests and type the three values printed by console.log.
const pattern = /colou?r/;
console.log(pattern.test("color"), pattern.test("colour"), pattern.test("colouur"));color and colour match, but colouur does not. The output is true true false.
Type the exact matched text.
const code = "Item 12345";
console.log(code.match(/\d{2,4}/)[0]);The first valid run has five digits, but {2,4} can take at most four, so it returns 1234.
Change the pattern so the match is only "one".
const text = 'say "one" and "two".';
console.log(text.match(/".+"/)[0]);text.match(/".+?"/)[0]
text.match(/"[^"]*"/)[0]Both return the first quoted string for this simple input. The negated class is often clearer when escaped quotes are not part of the format.
Predict the exact string printed by the replacement.
console.log("abc".replace(/x*/g, "-"));The replacement appears around every letter because the pattern can match an empty string at each position. It prints -a-b-c-.
Write a regex pattern for exactly four digits, such as a simple PIN. This exercise uses anchors; the anchors lesson will explain them in more detail.
/^\d{4}$/\d{4} counts four digits. ^ and $ make the check cover the whole string.
Check your understanding
8 QUESTIONSQuestion 1 of 8What does a quantifier repeat?
Choose an answer to see the explanation.
Question 2 of 8What does the empty-match replacement print?
Read the code, then predictconsole.log("abc".replace(/x*/g, "-"));Choose an answer to see the explanation.
Question 3 of 8What does the greedy quoted pattern print?
Read the code, then predictconst text = 'say "one" and "two".'; console.log(text.match(/".+"/)[0]);Choose an answer to see the explanation.
Question 4 of 8What does the lazy quoted pattern print?
Read the code, then predictconst text = 'say "one" and "two".'; console.log(text.match(/".+?"/)[0]);Choose an answer to see the explanation.
Question 5 of 8What does JavaScript do with
{,3}here?Read the code, then predictconsole.log(new RegExp("a{,3}").test("a{,3}")); try { new RegExp("a{,3}", "u"); } catch (error) { console.log(error.name); }Choose an answer to see the explanation.
Question 6 of 8Why does anchoring matter for length checks?
Read the code, then predictconsole.log(/\d{4}/.test("12345")); console.log(/^\d{4}$/.test("12345"));Choose an answer to see the explanation.
Question 7 of 8What does the quantified capture keep?
Read the code, then predictconst match = "abc".match(/([abc])+/); console.log(match[1]);Choose an answer to see the explanation.
Question 8 of 8Which pattern is the clearest simple quoted-string boundary when escapes are not part of the problem?
Choose an answer to see the explanation.
Key takeaways
- Quantifiers repeat the single preceding token; group tokens when you need a larger unit.
+needs at least one, while*and?can match zero.- Use
{n},{n,}, and{n,m}; use{0,m}, not{,m}. - Greedy quantifiers try the biggest span first and backtrack; lazy quantifiers try the smallest span first and expand.
- For simple quoted text, a negated class like
"[^"]*"often states the boundary more clearly than lazy dot.
Remember the one-liner.
Quantifiers control how many copies of one token can match; greedy and lazy control which legal size JavaScript tries first.
Up next: Groups & backreferences, where captures, non-capturing groups, alternation, and backreferences get their own full treatment.