cf.completefrontendCode editorOpen lab
THE JAVASCRIPT FIELD GUIDE

Scanning source text

Learn how JavaScript engines decode source bytes, scan UTF-16 text into tokens, handle identifiers, regex slashes, and optimize scanners.

By the end, you can
  • 01
    Trace bytes to tokensExplain how fetched or file bytes become decoded source text, UTF-16 positions, and a stream of lexical tokens.
  • 02
    Read tricky token boundariesPredict identifiers with Unicode escapes, contextual keywords, line terminators, comments, and slash-as-regex decisions.
  • 03
    Use scanner facts in real toolsCompare a small tokenizer with Acorn and connect scanner choices to highlighters, linters, minifiers, and engine performance.

Raw text becomes a token stream

Before JavaScript can build an AST, compile bytecode, or optimize a hot function, it has to answer a smaller question: where does one meaningful piece of source text end and the next one begin? The scanner, often called a lexer or tokenizer, performs that first pass over decoded source text.

Definition

A scanner reads decoded JavaScript source text and emits tokens: identifiers, keywords, punctuators, literals, and end markers with source ranges. The parser consumes tokens, not raw characters.

Real-life analogySplitting a sentence before grammar class

Imagine taking a sentence and first marking words, commas, and periods before you discuss subjects and verbs. Tokenizing is that first marking step. It does not yet understand the whole program; it prepares the pieces the parser can arrange.

In real life: Letters on the page
In JavaScript: Decoded source characters
In real life: Words and punctuation
In JavaScript: Tokens such as identifier, num, and =>
In real life: Grammar class diagramming the sentence
In JavaScript: The parser building an AST

Where the analogy stops: Human readers use meaning and context freely. A scanner follows exact language grammar and sometimes needs the parser to tell it which kind of slash is allowed.

This lesson builds on Code structure for automatic semicolon insertion, Unicode & strings for UTF-16, Text encoding & Base64 for bytes, and Regex basics for regex literals. It also points back to Meet the engines and ahead to Parsing & abstract syntax trees.

Source starts as bytes, not tokens

STEP THROUGH

A browser usually receives JavaScript as network bytes. A runtime reading from disk also starts with bytes. The host chooses a character encoding before the engine scans syntax: modules are decoded as UTF-8 by default, classic scripts still have legacy charset rules such as a document encoding or <script charset>, and a byte order mark can identify or be stripped from the beginning of the stream.

Bytes become decoded source text
Step 0 of 5Ready
Your turn: follow the blue line

Step through the path from bytes to the UTF-16 source text a scanner consumes.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
const decoded = new TextDecoder("utf-8").decode(bytes);const codeUnits = Array.from(  { length: decoded.length },  (_, index) => decoded.charCodeAt(index).toString(16));console.log(decoded);console.log(codeUnits.join(" "));
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

Once decoding finishes, the original byte offsets are no longer the scanner's main coordinate system. Parser locations, syntax errors, devtools ranges, and source maps usually talk in line/column pairs or UTF-16 offsets because that is how JavaScript strings index text.

Keep these representations separate
RepresentationUnitTiny exampleWhy it matters
Byte8-bit number from a file or response body0x63Before JavaScript syntax exists; the host decodes it.
Code pointUnicode catalog numberU+1F600Useful for Unicode properties such as ID_Start.
UTF-16 code unit16-bit JavaScript string storage unit0xD83D, 0xDE00Source positions and string indexes use these units.
TokenScanner result with one syntactic roleidentifier, num, /The parser consumes these instead of raw characters.

Engines scan a UTF-16 view

The V8 scanner article describes a UTF16CharacterStream abstraction. Chrome may hand V8 Latin-1, UTF-8, or UTF-16-backed data, and the stream presents a UTF-16 view to the scanner. The blog names the low-level method Utf16CharacterStream::Advance(): it returns the next UTF-16 code unit or -1 at end of input.

That does not mean every Unicode character is one unit. Astral characters use surrogate pairs, just as you saw in Unicode & strings. V8's post explains that the scanner avoids combining surrogate pairs until it needs Unicode identifier checks, because strings and comments can often be copied as code units.

Position rule of thumb

If a tool says token start: 12 and end: 14, read those as JavaScript string offsets: UTF-16 code units. A single visible symbol can cover two offsets when it is represented by a surrogate pair.

The scanner's daily work

STEP THROUGH

A scanner walks forward. At each cursor position it skips trivia, chooses the longest token that fits, records a source range, and hands the token to the parser. It recognizes identifiers, keywords, punctuators, numeric literals with separators or n, string escapes, template chunks, and regex literals when the grammar allows them.

Scan one line into tokens
Step 0 of 6Ready
Your turn: follow the blue line

Watch a small source line become keyword, identifier, punctuator, division, number, and semicolon tokens.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
const tokens = tokenizeSubset(source);console.log(tokens.map((token) => token.type + ":" + token.value).join(" | "));
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.
What a JavaScript scanner tracks
ConcernScanner responsibility
Whitespace and commentsSkip them, but remember whether a line terminator appeared for ASI.
Identifiers and keywordsRead Unicode identifier characters, cook escapes, then classify reserved words.
LiteralsCollect numbers, strings, template parts, and regular expressions with their flags.
PunctuatorsPrefer the longest valid punctuator such as >>>= before >>> or >.
LookaheadKeep enough lookahead for the current token and usually one token for the parser.

In classic browser scripts, Annex B also permits HTML-like comments such as a line starting with <!--. Treat that as legacy compatibility, not modern style, and do not rely on it in modules.

HTML-like comments are legacy script triviaJavaScript
<!-- valid as an HTML-like comment in classic scriptsconsole.log("after an Annex B comment");// Modules and non-browser tooling should not rely on this legacy form.
Which token bucket is this?
  • café
  • if
  • 1_000n
  • /ok/u
  • =>
  • /* line break\n */
Try it yourself
0 of 6 correct

Sort each source fragment into the bucket a scanner thinks about first.

Choose a category for every card. You can change an answer at any time; Reset clears them all.

Identifiers, keywords, and escapes

REAL ERRORS

Identifier starts use Unicode's ID_Start property, plus $ and _. Later characters use ID_Continue, which also admits digits and a few join controls. Escapes are cooked before the parser sees the name, so \u0061bc is the identifier abc.

Escaped identifiers and escaped keywordsPop out in the code editor (opens in a new tab)JavaScript
const \u0061bc = 1;console.log(abc); try {  new Function("const \\u0069f = 1;");} catch (error) {  console.log(error.name + ": " + error.message);}

The second half proves the engine-specific V8 message used by Node 22 and Chrome: SyntaxError: Keyword must not contain escaped characters. Acorn reports the same source as a syntax error too, with its own wording. Either way, an escape cannot hide a reserved keyword.

Some words are contextual. let, yield, await, async, of, get, set, and static gain special meaning only in particular grammar positions. The parser decides those contexts after the scanner has delivered the token text.

Why / is ambiguous: regex or division

ACORN + NODE

A slash can begin a regex literal, divide two expressions, start a comment, or combine with =. The scanner cannot decide from the slash alone. The ECMAScript lexical grammar has goal symbols such as InputElementDiv and InputElementRegExp; in practice, the parser drives the scanner by asking for the kind of input element that is valid in the current grammar position.

Slash puzzles in Node and ChromePop out in the code editor (opens in a new tab)JavaScript
let a, b = 20, c = 4, d = 2;a = b / c / d;console.log(a); try {  new Function('let a, b = true, d = "abc"; a = b\n/c/.test(d);');} catch (error) {  console.log(error.name + ": " + error.message);} try {  let x, y = 10;  x = y / 2 / g;} catch (error) {  console.log(error.name + ": " + error.message);}

Line 2 is ordinary division. The newline puzzle still fails as division because a = b can continue with a binary operator on the next line; V8 reports SyntaxError: Unexpected token '.'. The /g puzzle is also division, so runtime looks for a variable named g and throws ReferenceError.

Slash cases from the real tokenizers
Slash decisions4 cases
source"a = b / c / d;"
lessonidentifier:a@0-1 | punctuator:=@2-3 | identifier:b@4-5 | division:/@6-7 | identifier:c@8-9 | division:/@10-11 | identifier:d@12-13 | punctuator:;@13-14
acornname:a@0-1 | =:=@2-3 | name:b@4-5 | /:/@6-7 | name:c@8-9 | /:/@10-11 | name:d@12-13 | ;:;@13-14
source"value = /c/.test(name);"
lessonidentifier:value@0-5 | punctuator:=@6-7 | regexp:/c/@8-11 | punctuator:.@11-12 | identifier:test@12-16 | punctuator:(@16-17 | identifier:name@17-21 | punctuator:)@21-22 | punctuator:;@22-23
acornname:value@0-5 | =:=@6-7 | regexp:/c/@8-11 | .:.@11-12 | name:test@12-16 | (:(@16-17 | name:name@17-21 | ):)@21-22 | ;:;@22-23
source"a = b\n/c/.test(d);"
lessonidentifier:a@0-1 | punctuator:=@2-3 | identifier:b@4-5 | division:/@6-7↵ | identifier:c@7-8 | division:/@8-9 | punctuator:.@9-10 | identifier:test@10-14 | punctuator:(@14-15 | identifier:d@15-16 | punctuator:)@16-17 | punctuator:;@17-18
acornname:a@0-1 | =:=@2-3 | name:b@4-5 | /:/@6-7 | name:c@7-8 | /:/@8-9 | .:.@9-10 | name:test@10-14 | (:(@14-15 | name:d@15-16 | ):)@16-17 | ;:;@17-18
source"x = y / 2 /g;"
lessonidentifier:x@0-1 | punctuator:=@2-3 | identifier:y@4-5 | division:/@6-7 | number:2@8-9 | division:/@10-11 | identifier:g@11-12 | punctuator:;@12-13
acornname:x@0-1 | =:=@2-3 | name:y@4-5 | /:/@6-7 | num:2@8-9 | /:/@10-11 | name:g@11-12 | ;:;@12-13
Try it yourself
These cases are computed once from the lesson tokenizer and Acorn.

Look at the token right before /. An identifier or number usually means division; a position that expects a new expression can allow a regex literal.

The newline case is intentionally surprising: a = b can still continue with division on the next line, so /c/ is not a regex there.

Build and compare a small tokenizer

LIVE LAB

The lesson's tokenizer is intentionally small but real: it skips comments and whitespace, recognizes identifiers, keywords, numbers, BigInts, strings, templates, punctuators, and chooses regex versus division from the previous token. It is not a full JavaScript implementation; Acorn is here so you can see exactly where a production parser differs.

Tokenizer loop excerptJavaScript
function tokenizeSubset(source) {  // Real lesson code also handles comments, strings, numbers, identifiers,  // and a previous-token decision for slash. This excerpt shows the loop shape.  const tokens = [];  let cursor = 0;  let previous = null;   while (cursor < source.length) {    const next = scanOneToken(source, cursor, previous);    if (next) {      tokens.push(next.token);      previous = next.token;      cursor = next.end;    } else {      cursor += 1;    }  }   return tokens;}
Compare the lesson tokenizer with Acorn
Editable source snapshotJavaScript
let café = 1_000n;const ok = /ok/u.test(name);x = y / 2 /g;
Token streams23 lesson tokens

Lesson tokenizer

typevaluerange
contextuallet0–3
identifiercafé4–8
punctuator=9–10
bigint1_000n11–17
punctuator;17–18
keywordconst — line terminator before this token19–24
identifierok25–27
punctuator=28–29
regexp/ok/u30–35
punctuator.35–36
identifiertest36–40
punctuator(40–41
identifiername41–45
punctuator)45–46
punctuator;46–47
identifierx — line terminator before this token48–49
punctuator=50–51
identifiery52–53
division/54–55
number256–57
division/58–59
identifierg59–60
punctuator;60–61

Acorn tokenizer

typevaluerange
namelet0–3
namecafé4–8
==9–10
num1000n11–17
;;17–18
constconst19–24
nameok25–27
==28–29
regexp/ok/u30–35
..35–36
nametest36–40
((40–41
namename41–45
))45–46
;;46–47
namex48–49
==50–51
namey52–53
//54–55
num256–57
//58–59
nameg59–60
;;60–61

Differences to inspect

  • No boundary differences for this input.
Try it yourself

For this subset, the lesson tokenizer and Acorn agree on token boundaries.

Acorn is the same declared parser dependency used in other CompleteFrontend lessons. It provides the real token labels, values, and UTF-16 ranges.

Scanner performance is parser performance

Scanning happens before code can run, so engines spend real effort making it fast. V8 reports that its scanner improvements made single-token scanning roughly 1.4× faster, string scanning 1.3× faster, multiline comment scanning 2.1× faster, and identifier scanning 1.2–1.5× faster depending on identifier length.

Performance techniques from the V8 scanner article
TechniqueWhy it helps
ASCII fast pathsMost production source is ASCII, so V8 uses compact tables for identifier-start and identifier-continue checks.
Keyword perfect hashingV8's blog says it uses gperf to compute a perfect hash from keyword length and first characters.
Whitespace loopsWhitespace and comments are skipped quickly while preserving the line-break fact ASI needs.
AdvanceUntilV8 added an interface that lets scan helpers consume many code units with less per-character overhead.

The developer-facing advice is modest: write clear code, let your build pipeline minify production bundles, and avoid non-ASCII identifiers when there is no product reason for them. The engine work is about startup cost and tooling scale, not micro-optimizing one variable name by hand.

Where this shows up in real tools

Syntax highlighters, formatters, linters, minifiers, bundlers, source-map tools, and browser error messages all depend on token boundaries. A highlighter wants to know that /ok/u is a regex literal; a minifier wants to remove comments without accidentally changing ASI; a linter wants to report the exact range of an escaped keyword error.

Many token kinds in one short programPop out in the code editor (opens in a new tab)JavaScript
const total = 1_000n;const word = "token";const okay = /tok(en)?/u.test(word);console.log(total, word, okay);

The next lesson, Parsing & abstract syntax trees, takes these tokens and builds nested structure: declarations, expressions, statements, and early errors.

Common misconceptions

“The scanner reads bytes.”

The host decodes bytes first; engines scan a character stream and report UTF-16 source positions.

“`length` equals characters.”

JavaScript source positions are UTF-16 code units, so astral code points take two positions.

“Keywords are just identifiers.”

They are identifier-shaped text, but scanners return distinct tokens for reserved keywords.

“A slash always starts a regex.”

After an expression, slash is division; after a token that expects an expression, slash can start a regex literal.

“Comments vanish completely.”

Their text is skipped, but line terminators inside them can still affect ASI.

Lookalikes at scanner time
Looks likeActuallySafe mental model
/x/Regex only in a regex-allowed grammar goalAsk what previous token allows next.
\u0061bcThe identifier name abcEscapes are cooked before parsing names.
😀 in sourceTwo UTF-16 code unitsTool ranges use JavaScript string offsets.
/* comment */Skipped triviaSkipped does not mean line breaks are forgotten.

Practice exercises

5 EXERCISES
Exercise 1 · Warm-upCook an identifier escape

Type the identifier name that the escape creates.

Starter codePop out in the code editor (opens in a new tab)JavaScript
const \u0061bc = 7;
console.log(abc);

Answer, then press Check. Spacing and letter case don’t matter.

    Exercise 2 · PracticeLabel the escaped keyword error

    Run the message mentally and type the error name printed by the catch block.

    Starter codePop out in the code editor (opens in a new tab)JavaScript
    try {
      new Function("const \\u0069f = 1;");
    } catch (error) {
      console.log(error.name + ": " + error.message);
    }

    Answer, then press Check. Spacing and letter case don’t matter.

      Exercise 3 · PracticeEvaluate two division tokens

      Both slashes are division. Type the number that prints.

      Starter codePop out in the code editor (opens in a new tab)JavaScript
      let a, b = 20, c = 4, d = 2;
      a = b / c / d;
      console.log(a);

      Answer, then press Check. Spacing and letter case don’t matter.

        Exercise 4 · PracticeRecognize a regex literal

        Type the text printed when the regex test succeeds.

        Starter codePop out in the code editor (opens in a new tab)JavaScript
        const name = "café";
        if (/caf\u00e9/u.test(name)) {
          console.log("match");
        }

        Answer, then press Check. Spacing and letter case don’t matter.

          Exercise 5 · ChallengeExplain the /g surprise

          Type the error name printed by the catch block.

          Starter codePop out in the code editor (opens in a new tab)JavaScript
          try {
            let x, y = 10;
            x = y / 2 / g;
          } catch (error) {
            console.log(error.name);
          }

          Answer, then press Check. Spacing and letter case don’t matter.

            Check your understanding

            8 QUESTIONS
            Lesson quiz · 8 questionsScore: first tries count
            1. Question 1 of 8What is the scanner's main job?

              Choose an answer to see the explanation.

            2. Question 2 of 8What does the escaped identifier snippet print?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              const \u0061bc = 3;
              console.log(abc);

              Choose an answer to see the explanation.

            3. Question 3 of 8What does V8/Node 22 report for an escaped keyword spelling?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              try {
                new Function("const \\u0069f = 1;");
              } catch (error) {
                console.log(error.name + ": " + error.message);
              }

              Choose an answer to see the explanation.

            4. Question 4 of 8Why does the scanner track line terminators while skipping whitespace and comments?

              Choose an answer to see the explanation.

            5. Question 5 of 8What does the division snippet print?

              Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
              let a, b = 20, c = 4, d = 2;
              a = b / c / d;
              console.log(a);

              Choose an answer to see the explanation.

            6. Question 6 of 8Which grammar fact explains regex versus division slash?

              Choose an answer to see the explanation.

            7. Question 7 of 8Which item is a contextual keyword rather than always a reserved keyword?

              Choose an answer to see the explanation.

            8. Question 8 of 8Which V8 scanner optimization claim is from the V8 blog?

              Choose an answer to see the explanation.

            Key takeaways

            • Bytes are decoded before scanning; JavaScript source positions are based on decoded text, usually UTF-16 code units.
            • The scanner skips trivia, recognizes the longest valid token, and keeps line-break facts for ASI.
            • Identifier escapes are cooked, but escaped keywords still produce syntax errors.
            • Slash is decided by grammar context: regex in regex-allowed positions, division after completed expressions.
            • Scanner performance matters because no parser, compiler, or optimizer can run before tokenization starts.

            Next, Parsing & abstract syntax trees shows how these tokens become a tree the engine can compile.

            CompleteFrontend Clear concepts. Working examples.