Scanning source text
Learn how JavaScript engines decode source bytes, scan UTF-16 text into tokens, handle identifiers, regex slashes, and optimize scanners.
- 01Trace bytes to tokensExplain how fetched or file bytes become decoded source text, UTF-16 positions, and a stream of lexical tokens.
- 02Read tricky token boundariesPredict identifiers with Unicode escapes, contextual keywords, line terminators, comments, and slash-as-regex decisions.
- 03Use scanner facts in real toolsCompare a small tokenizer with Acorn and connect scanner choices to highlighters, linters, minifiers, and engine performance.
Raw text becomes a token stream
Before JavaScript can build an AST, compile bytecode, or optimize a hot function, it has to answer a smaller question: where does one meaningful piece of source text end and the next one begin? The scanner, often called a lexer or tokenizer, performs that first pass over decoded source text.
A scanner reads decoded JavaScript source text and emits tokens: identifiers, keywords, punctuators, literals, and end markers with source ranges. The parser consumes tokens, not raw characters.
Imagine taking a sentence and first marking words, commas, and periods before you discuss subjects and verbs. Tokenizing is that first marking step. It does not yet understand the whole program; it prepares the pieces the parser can arrange.
- In real life: Letters on the page
- In JavaScript: Decoded source characters
- In real life: Words and punctuation
- In JavaScript: Tokens such as
identifier,num, and=> - In real life: Grammar class diagramming the sentence
- In JavaScript: The parser building an AST
Where the analogy stops: Human readers use meaning and context freely. A scanner follows exact language grammar and sometimes needs the parser to tell it which kind of slash is allowed.
This lesson builds on Code structure for automatic semicolon insertion, Unicode & strings for UTF-16, Text encoding & Base64 for bytes, and Regex basics for regex literals. It also points back to Meet the engines and ahead to Parsing & abstract syntax trees.
Source starts as bytes, not tokens
STEP THROUGHA browser usually receives JavaScript as network bytes. A runtime reading from disk also starts with bytes. The host chooses a character encoding before the engine scans syntax: modules are decoded as UTF-8 by default, classic scripts still have legacy charset rules such as a document encoding or <script charset>, and a byte order mark can identify or be stripped from the beginning of the stream.
Step through the path from bytes to the UTF-16 source text a scanner consumes.
script
const decoded = new TextDecoder("utf-8").decode(bytes);const codeUnits = Array.from( { length: decoded.length }, (_, index) => decoded.charCodeAt(index).toString(16));console.log(decoded);console.log(codeUnits.join(" "));Once decoding finishes, the original byte offsets are no longer the scanner's main coordinate system. Parser locations, syntax errors, devtools ranges, and source maps usually talk in line/column pairs or UTF-16 offsets because that is how JavaScript strings index text.
| Representation | Unit | Tiny example | Why it matters |
|---|---|---|---|
| Byte | 8-bit number from a file or response body | 0x63 | Before JavaScript syntax exists; the host decodes it. |
| Code point | Unicode catalog number | U+1F600 | Useful for Unicode properties such as ID_Start. |
| UTF-16 code unit | 16-bit JavaScript string storage unit | 0xD83D, 0xDE00 | Source positions and string indexes use these units. |
| Token | Scanner result with one syntactic role | identifier, num, / | The parser consumes these instead of raw characters. |
Engines scan a UTF-16 view
The V8 scanner article describes a UTF16CharacterStream abstraction. Chrome may hand V8 Latin-1, UTF-8, or UTF-16-backed data, and the stream presents a UTF-16 view to the scanner. The blog names the low-level method Utf16CharacterStream::Advance(): it returns the next UTF-16 code unit or -1 at end of input.
That does not mean every Unicode character is one unit. Astral characters use surrogate pairs, just as you saw in Unicode & strings. V8's post explains that the scanner avoids combining surrogate pairs until it needs Unicode identifier checks, because strings and comments can often be copied as code units.
If a tool says token start: 12 and end: 14, read those as JavaScript string offsets: UTF-16 code units. A single visible symbol can cover two offsets when it is represented by a surrogate pair.
The scanner's daily work
STEP THROUGHA scanner walks forward. At each cursor position it skips trivia, chooses the longest token that fits, records a source range, and hands the token to the parser. It recognizes identifiers, keywords, punctuators, numeric literals with separators or n, string escapes, template chunks, and regex literals when the grammar allows them.
Watch a small source line become keyword, identifier, punctuator, division, number, and semicolon tokens.
script
const tokens = tokenizeSubset(source);console.log(tokens.map((token) => token.type + ":" + token.value).join(" | "));| Concern | Scanner responsibility |
|---|---|
| Whitespace and comments | Skip them, but remember whether a line terminator appeared for ASI. |
| Identifiers and keywords | Read Unicode identifier characters, cook escapes, then classify reserved words. |
| Literals | Collect numbers, strings, template parts, and regular expressions with their flags. |
| Punctuators | Prefer the longest valid punctuator such as >>>= before >>> or >. |
| Lookahead | Keep enough lookahead for the current token and usually one token for the parser. |
In classic browser scripts, Annex B also permits HTML-like comments such as a line starting with <!--. Treat that as legacy compatibility, not modern style, and do not rely on it in modules.
<!-- valid as an HTML-like comment in classic scriptsconsole.log("after an Annex B comment");// Modules and non-browser tooling should not rely on this legacy form.caféif1_000n/ok/u=>/* line break\n */
Sort each source fragment into the bucket a scanner thinks about first.
Identifiers, keywords, and escapes
REAL ERRORSIdentifier starts use Unicode's ID_Start property, plus $ and _. Later characters use ID_Continue, which also admits digits and a few join controls. Escapes are cooked before the parser sees the name, so \u0061bc is the identifier abc.
const \u0061bc = 1;console.log(abc); try { new Function("const \\u0069f = 1;");} catch (error) { console.log(error.name + ": " + error.message);}The second half proves the engine-specific V8 message used by Node 22 and Chrome: SyntaxError: Keyword must not contain escaped characters. Acorn reports the same source as a syntax error too, with its own wording. Either way, an escape cannot hide a reserved keyword.
Some words are contextual. let, yield, await, async, of, get, set, and static gain special meaning only in particular grammar positions. The parser decides those contexts after the scanner has delivered the token text.
Why / is ambiguous: regex or division
ACORN + NODEA slash can begin a regex literal, divide two expressions, start a comment, or combine with =. The scanner cannot decide from the slash alone. The ECMAScript lexical grammar has goal symbols such as InputElementDiv and InputElementRegExp; in practice, the parser drives the scanner by asking for the kind of input element that is valid in the current grammar position.
let a, b = 20, c = 4, d = 2;a = b / c / d;console.log(a); try { new Function('let a, b = true, d = "abc"; a = b\n/c/.test(d);');} catch (error) { console.log(error.name + ": " + error.message);} try { let x, y = 10; x = y / 2 / g;} catch (error) { console.log(error.name + ": " + error.message);}Line 2 is ordinary division. The newline puzzle still fails as division because a = b can continue with a binary operator on the next line; V8 reports SyntaxError: Unexpected token '.'. The /g puzzle is also division, so runtime looks for a variable named g and throws ReferenceError.
"a = b / c / d;"identifier:a@0-1 | punctuator:=@2-3 | identifier:b@4-5 | division:/@6-7 | identifier:c@8-9 | division:/@10-11 | identifier:d@12-13 | punctuator:;@13-14name:a@0-1 | =:=@2-3 | name:b@4-5 | /:/@6-7 | name:c@8-9 | /:/@10-11 | name:d@12-13 | ;:;@13-14"value = /c/.test(name);"identifier:value@0-5 | punctuator:=@6-7 | regexp:/c/@8-11 | punctuator:.@11-12 | identifier:test@12-16 | punctuator:(@16-17 | identifier:name@17-21 | punctuator:)@21-22 | punctuator:;@22-23name:value@0-5 | =:=@6-7 | regexp:/c/@8-11 | .:.@11-12 | name:test@12-16 | (:(@16-17 | name:name@17-21 | ):)@21-22 | ;:;@22-23"a = b\n/c/.test(d);"identifier:a@0-1 | punctuator:=@2-3 | identifier:b@4-5 | division:/@6-7↵ | identifier:c@7-8 | division:/@8-9 | punctuator:.@9-10 | identifier:test@10-14 | punctuator:(@14-15 | identifier:d@15-16 | punctuator:)@16-17 | punctuator:;@17-18name:a@0-1 | =:=@2-3 | name:b@4-5 | /:/@6-7 | name:c@7-8 | /:/@8-9 | .:.@9-10 | name:test@10-14 | (:(@14-15 | name:d@15-16 | ):)@16-17 | ;:;@17-18"x = y / 2 /g;"identifier:x@0-1 | punctuator:=@2-3 | identifier:y@4-5 | division:/@6-7 | number:2@8-9 | division:/@10-11 | identifier:g@11-12 | punctuator:;@12-13name:x@0-1 | =:=@2-3 | name:y@4-5 | /:/@6-7 | num:2@8-9 | /:/@10-11 | name:g@11-12 | ;:;@12-13Look at the token right before /. An identifier or number usually means division; a position that expects a new expression can allow a regex literal.
a = b can still continue with division on the next line, so /c/ is not a regex there.Build and compare a small tokenizer
LIVE LABThe lesson's tokenizer is intentionally small but real: it skips comments and whitespace, recognizes identifiers, keywords, numbers, BigInts, strings, templates, punctuators, and chooses regex versus division from the previous token. It is not a full JavaScript implementation; Acorn is here so you can see exactly where a production parser differs.
function tokenizeSubset(source) { // Real lesson code also handles comments, strings, numbers, identifiers, // and a previous-token decision for slash. This excerpt shows the loop shape. const tokens = []; let cursor = 0; let previous = null; while (cursor < source.length) { const next = scanOneToken(source, cursor, previous); if (next) { tokens.push(next.token); previous = next.token; cursor = next.end; } else { cursor += 1; } } return tokens;}let café = 1_000n;const ok = /ok/u.test(name);x = y / 2 /g;Lesson tokenizer
| type | value | range |
|---|---|---|
contextual | let | 0–3 |
identifier | café | 4–8 |
punctuator | = | 9–10 |
bigint | 1_000n | 11–17 |
punctuator | ; | 17–18 |
keyword | const — line terminator before this token | 19–24 |
identifier | ok | 25–27 |
punctuator | = | 28–29 |
regexp | /ok/u | 30–35 |
punctuator | . | 35–36 |
identifier | test | 36–40 |
punctuator | ( | 40–41 |
identifier | name | 41–45 |
punctuator | ) | 45–46 |
punctuator | ; | 46–47 |
identifier | x — line terminator before this token | 48–49 |
punctuator | = | 50–51 |
identifier | y | 52–53 |
division | / | 54–55 |
number | 2 | 56–57 |
division | / | 58–59 |
identifier | g | 59–60 |
punctuator | ; | 60–61 |
Acorn tokenizer
| type | value | range |
|---|---|---|
name | let | 0–3 |
name | café | 4–8 |
= | = | 9–10 |
num | 1000n | 11–17 |
; | ; | 17–18 |
const | const | 19–24 |
name | ok | 25–27 |
= | = | 28–29 |
regexp | /ok/u | 30–35 |
. | . | 35–36 |
name | test | 36–40 |
( | ( | 40–41 |
name | name | 41–45 |
) | ) | 45–46 |
; | ; | 46–47 |
name | x | 48–49 |
= | = | 50–51 |
name | y | 52–53 |
/ | / | 54–55 |
num | 2 | 56–57 |
/ | / | 58–59 |
name | g | 59–60 |
; | ; | 60–61 |
Differences to inspect
- No boundary differences for this input.
For this subset, the lesson tokenizer and Acorn agree on token boundaries.
Scanner performance is parser performance
Scanning happens before code can run, so engines spend real effort making it fast. V8 reports that its scanner improvements made single-token scanning roughly 1.4× faster, string scanning 1.3× faster, multiline comment scanning 2.1× faster, and identifier scanning 1.2–1.5× faster depending on identifier length.
| Technique | Why it helps |
|---|---|
| ASCII fast paths | Most production source is ASCII, so V8 uses compact tables for identifier-start and identifier-continue checks. |
| Keyword perfect hashing | V8's blog says it uses gperf to compute a perfect hash from keyword length and first characters. |
| Whitespace loops | Whitespace and comments are skipped quickly while preserving the line-break fact ASI needs. |
| AdvanceUntil | V8 added an interface that lets scan helpers consume many code units with less per-character overhead. |
The developer-facing advice is modest: write clear code, let your build pipeline minify production bundles, and avoid non-ASCII identifiers when there is no product reason for them. The engine work is about startup cost and tooling scale, not micro-optimizing one variable name by hand.
Where this shows up in real tools
Syntax highlighters, formatters, linters, minifiers, bundlers, source-map tools, and browser error messages all depend on token boundaries. A highlighter wants to know that /ok/u is a regex literal; a minifier wants to remove comments without accidentally changing ASI; a linter wants to report the exact range of an escaped keyword error.
const total = 1_000n;const word = "token";const okay = /tok(en)?/u.test(word);console.log(total, word, okay);The next lesson, Parsing & abstract syntax trees, takes these tokens and builds nested structure: declarations, expressions, statements, and early errors.
Common misconceptions
“The scanner reads bytes.”
The host decodes bytes first; engines scan a character stream and report UTF-16 source positions.
“`length` equals characters.”
JavaScript source positions are UTF-16 code units, so astral code points take two positions.
“Keywords are just identifiers.”
They are identifier-shaped text, but scanners return distinct tokens for reserved keywords.
“A slash always starts a regex.”
After an expression, slash is division; after a token that expects an expression, slash can start a regex literal.
“Comments vanish completely.”
Their text is skipped, but line terminators inside them can still affect ASI.
| Looks like | Actually | Safe mental model |
|---|---|---|
/x/ | Regex only in a regex-allowed grammar goal | Ask what previous token allows next. |
\u0061bc | The identifier name abc | Escapes are cooked before parsing names. |
😀 in source | Two UTF-16 code units | Tool ranges use JavaScript string offsets. |
/* comment */ | Skipped trivia | Skipped does not mean line breaks are forgotten. |
Practice exercises
5 EXERCISESType the identifier name that the escape creates.
const \u0061bc = 7;
console.log(abc);\u0061bc cooks to abc, so the program prints the value stored in abc.
Run the message mentally and type the error name printed by the catch block.
try {
new Function("const \\u0069f = 1;");
} catch (error) {
console.log(error.name + ": " + error.message);
}V8 throws a SyntaxError before the generated function is created.
Both slashes are division. Type the number that prints.
let a, b = 20, c = 4, d = 2;
a = b / c / d;
console.log(a);20 / 4 is 5, then 5 / 2 is 2.5.
Type the text printed when the regex test succeeds.
const name = "café";
if (/caf\u00e9/u.test(name)) {
console.log("match");
}The regex matches café, so the block logs match.
/g surpriseType the error name printed by the catch block.
try {
let x, y = 10;
x = y / 2 / g;
} catch (error) {
console.log(error.name);
}The program parses as x = y / 2 / g; because no variable named g exists, runtime throws ReferenceError.
Check your understanding
8 QUESTIONSQuestion 1 of 8What is the scanner's main job?
Choose an answer to see the explanation.
Question 2 of 8What does the escaped identifier snippet print?
Read the code, then predictconst \u0061bc = 3; console.log(abc);Choose an answer to see the explanation.
Question 3 of 8What does V8/Node 22 report for an escaped keyword spelling?
Read the code, then predicttry { new Function("const \\u0069f = 1;"); } catch (error) { console.log(error.name + ": " + error.message); }Choose an answer to see the explanation.
Question 4 of 8Why does the scanner track line terminators while skipping whitespace and comments?
Choose an answer to see the explanation.
Question 5 of 8What does the division snippet print?
Read the code, then predictlet a, b = 20, c = 4, d = 2; a = b / c / d; console.log(a);Choose an answer to see the explanation.
Question 6 of 8Which grammar fact explains regex versus division slash?
Choose an answer to see the explanation.
Question 7 of 8Which item is a contextual keyword rather than always a reserved keyword?
Choose an answer to see the explanation.
Question 8 of 8Which V8 scanner optimization claim is from the V8 blog?
Choose an answer to see the explanation.
Key takeaways
- Bytes are decoded before scanning; JavaScript source positions are based on decoded text, usually UTF-16 code units.
- The scanner skips trivia, recognizes the longest valid token, and keeps line-break facts for ASI.
- Identifier escapes are cooked, but escaped keywords still produce syntax errors.
- Slash is decided by grammar context: regex in regex-allowed positions, division after completed expressions.
- Scanner performance matters because no parser, compiler, or optimizer can run before tokenization starts.
Next, Parsing & abstract syntax trees shows how these tokens become a tree the engine can compile.