Parsing & abstract syntax trees
Learn how JavaScript parsers turn tokens into ESTree-shaped ASTs, report early errors, handle cover grammars, and power tools.
- 01Trace parsingExplain how recursive-descent and Pratt parsers consume tokens into nested expression nodes.
- 02Read ASTsRecognize ESTree node types and compare a small parser's output with Acorn's real AST shape.
- 03Separate parser failuresTell syntax, early errors, cover-grammar reparsing, and runtime errors apart in real code.
The parser's job
The scanner turns source text into tokens. The parser consumes those tokens, checks the grammar, reports syntax and early errors, and builds a tree. Later phases can compile that tree to bytecode, while tools can inspect a similar tree to lint, format, transform, or bundle your code.
Parsing is the grammar-checking step that turns a token stream into a syntax tree. An abstract syntax tree, or AST, keeps the useful program structure while leaving out many concrete punctuation details.
This lesson builds on code structure, the precedence table from other operators, the Acorn module graph in bundlers, and the AST transform pipeline in transpilers & polyfills. We will also connect to linting, destructuring, arrow functions, and strict mode.
In grammar class, a sentence diagram shows how words attach to one another. A parser does that for JavaScript, except the rules are formal and every token must fit.
- In real life: Words in a sentence
- In JavaScript: Scanner tokens such as identifiers, numbers, and operators
- In real life: Grammar rules
- In JavaScript: ECMAScript productions such as
ExpressionandStatement - In real life: A sentence diagram
- In JavaScript: A tree that shows which parts belong together
- In real life: A simplified role diagram
- In JavaScript: An AST that keeps meaning but drops extra punctuation nodes
Where the analogy stops: Human language tolerates ambiguity and judgment. JavaScript parsing must choose one exact structure or report an error before code runs.
| Artifact | What it is | Example | Where it appears |
|---|---|---|---|
| Token | A classified chunk from the scanner | identifier(a), +, number(2) | Parser input |
| Concrete syntax tree / parse tree | A tree that mirrors grammar productions and punctuation | Parentheses, commas, and every grammar layer | Specs and parser internals |
| Abstract syntax tree | A smaller semantic tree for tools and compilers | BinaryExpression, CallExpression | Linters, compilers, formatters |
| Bytecode | Executable instructions for an interpreter | V8 Ignition bytecode | Engine execution after parsing |
Recursive-descent parsing
REAL PARSERA recursive-descent parser is organized as functions that match grammar rules. A parser for declarations calls a parser for statements. A parser for statements calls a parser for expressions. Parentheses, blocks, function bodies, and nested calls naturally become recursive calls.
| Technique | Mental model | Best fit |
|---|---|---|
| Recursive descent | One function per grammar rule | Statements, declarations, blocks, primary expressions |
| Pratt / precedence climbing | A loop guided by binding power | Binary expressions, ** right associativity, member/call postfixes |
| Cover grammar | Parse a broad shape first, reinterpret later | (a, b) as either expression grouping or arrow parameters |
Expressions need one extra trick: precedence. The mini parser in this lesson uses a Pratt parser, also called precedence climbing. Each binary operator has a binding power. The parser reads a left operand, then keeps consuming operators whose binding power is high enough for the current context.
Step through the real mini parser on a + b * c(d). Watch the call stack and the left or right node change as Pratt binding power groups the expression.
script
const input = "a + b * c(d)";const tokens = tokenize(input);let cursor = 0; function parseExpression(minBP = 0) { let left = parsePostfix(); while (isBinary(peek()) && bindingPower(peek()) >= minBP) { const operator = consume(); const right = parseExpression(nextMinBP(operator)); left = binary(left, operator, right); } return left;} function parsePostfix() { let expression = parsePrimary(); while (peek() === "(" || peek() === ".") { expression = parseCallOrMember(expression); } return expression;} function parsePrimary() { return tokenToNode(consume());} function tokenize(source) { return source.match(/[A-Za-z_$][A-Za-z0-9_$]*|[0-9]+|[*][*]|[()+*/.,-]/g).concat("<eof>");}function peek() { return tokens[cursor]; }function consume() { return tokens[cursor++]; }function isBinary(token) { return ["+", "-", "*", "/", "**"].includes(token); }function bindingPower(token) { return token === "+" || token === "-" ? 10 : token === "*" || token === "/" ? 20 : token === "**" ? 30 : -1;}function nextMinBP(operator) { return operator === "**" ? bindingPower(operator) : bindingPower(operator) + 1; }function tokenToNode(token) { return /^[A-Za-z_$]/.test(token) ? { type: "Identifier", name: token } : { type: "Literal", value: Number(token), raw: token }; }function parseCallOrMember(expression) { if (peek() === ".") return { type: "MemberExpression", object: expression, property: tokenToNode((consume(), consume())) }; consume(); const args = peek() === ")" ? [] : [parseExpression(0)]; consume(); return { type: "CallExpression", callee: expression, arguments: args };}function binary(left, operator, right) { return { type: "BinaryExpression", left, operator, right }; } Step through line by line. The source has the same shape a production parser uses: one function reads expressions, one reads postfix operators such as calls and member access, and one reads primary tokens. The recording is from the lesson's real parser on a + b * c(d).
Step through a second run of the same parser. This time the tree grows in the postfix loop, proving calls and member access bind before binary operators.
script
const input = "a + b * c(d)";const tokens = tokenize(input);let cursor = 0; function parseExpression(minBP = 0) { let left = parsePostfix(); while (isBinary(peek()) && bindingPower(peek()) >= minBP) { const operator = consume(); const right = parseExpression(nextMinBP(operator)); left = binary(left, operator, right); } return left;} function parsePostfix() { let expression = parsePrimary(); while (peek() === "(" || peek() === ".") { expression = parseCallOrMember(expression); } return expression;} function parsePrimary() { return tokenToNode(consume());} function tokenize(source) { return source.match(/[A-Za-z_$][A-Za-z0-9_$]*|[0-9]+|[*][*]|[()+*/.,-]/g).concat("<eof>");}function peek() { return tokens[cursor]; }function consume() { return tokens[cursor++]; }function isBinary(token) { return ["+", "-", "*", "/", "**"].includes(token); }function bindingPower(token) { return token === "+" || token === "-" ? 10 : token === "*" || token === "/" ? 20 : token === "**" ? 30 : -1;}function nextMinBP(operator) { return operator === "**" ? bindingPower(operator) : bindingPower(operator) + 1; }function tokenToNode(token) { return /^[A-Za-z_$]/.test(token) ? { type: "Identifier", name: token } : { type: "Literal", value: Number(token), raw: token }; }function parseCallOrMember(expression) { if (peek() === ".") return { type: "MemberExpression", object: expression, property: tokenToNode((consume(), consume())) }; consume(); const args = peek() === ")" ? [] : [parseExpression(0)]; consume(); return { type: "CallExpression", callee: expression, arguments: args };}function binary(left, operator, right) { return { type: "BinaryExpression", left, operator, right }; } The second replay proves why config.get(user.name).theme is one chain before binary operators get a chance to group anything. Calls and member access are postfix operations: they attach directly to the expression already built.
Abstract syntax trees and ESTree
NODE SHAPESA concrete syntax tree, or parse tree, mirrors grammar productions. It can include layers for parentheses, comma lists, and punctuation that matter while parsing but are noisy for most tools. An AST is more abstract: it keeps the meaning-bearing nodes.
ESTree is the community AST shape used by many JavaScript tools. Acorn, Espree, AST Explorer, ESLint rules, and many Babel-compatible workflows expose ESTree or ESTree-like nodes. V8 does build an AST internally, and the V8 scanner post says that AST is compiled to Ignition bytecode, but V8's internal node classes are not ESTree.
| Node | Represents | Important fields |
|---|---|---|
Program | The whole script or module | body list, sourceType |
VariableDeclaration | let, const, or var declarations | kind, declarations |
BinaryExpression | An infix operator expression | left, operator, right |
CallExpression | A function or method call | callee, arguments, optional |
MemberExpression | Property access such as obj.name | object, property, computed |
ArrowFunctionExpression | Arrow functions after cover-grammar checks | params, body, expression |
{
"type": "Program",
"body": [
{
"type": "ExpressionStatement",
"expression": {
"type": "BinaryExpression",
"left": {
"type": "Identifier",
"name": "a"
},
"operator": "+",
"right": {
"type": "BinaryExpression",
"left": {
"type": "Identifier",
"name": "b"
},
"operator": "*",
"right": {
"type": "CallExpression",
"callee": {
"type": "Identifier",
"name": "c"
},
"arguments": [
{
"type": "Identifier",
"name": "d"
}
],
"optional": false
}
}
}
}
],
"sourceType": "script"
}{
"type": "Program",
"body": [
{
"type": "ExpressionStatement",
"expression": {
"type": "BinaryExpression",
"left": {
"type": "Identifier",
"name": "a"
},
"operator": "+",
"right": {
"type": "BinaryExpression",
"left": {
"type": "Identifier",
"name": "b"
},
"operator": "*",
"right": {
"type": "CallExpression",
"callee": {
"type": "Identifier",
"name": "c"
},
"arguments": [
{
"type": "Identifier",
"name": "d"
}
],
"optional": false
}
}
}
}
],
"sourceType": "script"
}The stripped trees match for this supported subset.
V8's scanner article describes parsing source into an AST and compiling that AST to Ignition bytecode. The ESTree project describes a community format that caught on as a lingua franca for JavaScript tools. Those are related ideas, not the same representation.
Precedence, associativity, and one sharp SyntaxError
REAL OUTPUTThe operator lesson introduced precedence. A parser turns that table into structure. Multiplication nests inside addition, member access nests before multiplication, and exponentiation is right-associative.
** and the unary-left SyntaxErrorconsole.log(2 ** 3 ** 2); try { new Function("return -2 ** 2;"); console.log("parsed");} catch (error) { console.log(error.name + ": " + error.message);}Line 1 prints 512, because JavaScript groups 2 ** 3 ** 2 as 2 ** (3 ** 2). Line 4 proves the special grammar restriction: a unary expression cannot appear directly on the left of **. Use -(2 ** 2) or (-2) ** 2 when you mean one of those shapes.
Early errors stop the whole script
PARSE VS RUNThe ECMAScript specification defines early errors: errors detected and reported before evaluating the script, module, or eval code that contains them. This is stronger than “the engine could guess it will fail.” The spec says which static checks are early errors.
const early = [];try { eval('early.push("before"); let item = 1; let item = 2; early.push("after");');} catch (error) { early.push(error.name);}console.log(early.join(" -> ")); const runtime = [];try { eval('runtime.push("before"); missingFunction(); runtime.push("after");');} catch (error) { runtime.push(error.name);}console.log(runtime.join(" -> "));The first eval string has a duplicate let. It records only SyntaxError; the earlier early.push("before") never runs. The second eval string parses, pushes before, then throws when execution reaches missingFunction().
// These are fragments for discussion, not code to run together.let total = 1;let total = 2; // duplicate lexical declaration return 1; // outside a function1 = 2; // invalid assignment target "use strict";010; // legacy octal literal in strict codeconst cases = [ ["duplicate let", 'let count; let count;'], ["return outside function", 'return 1;'], ["invalid assignment target", '1 = 2;'], ["strict octal literal", '"use strict"; 010;'], ["strict duplicate parameters", '"use strict"; function f(a, a) {}'], ["await outside async", 'await 1;'],]; for (const [label, source] of cases) { try { eval(source); console.log(label + ": ran"); } catch (error) { console.log(label + ": " + error.name); }}`let count; let count;``console.log('before'); missingFunction();``1 = 2;``'use strict'; 010;``console.log('before'); throw new Error('boom');``const add = (a, b) => a + b;`
Sort each snippet by when JavaScript reports the problem. Remember: early errors stop all evaluation of that script or eval string.
Cover grammars and reparsing arrow parameters
AMBIGUITYSome JavaScript token sequences are ambiguous until a later token appears. The spec uses cover grammars: broad grammar shapes that temporarily cover more than one final meaning. The classic example is (a, b). Without =>, it is a parenthesized comma expression. With =>, it becomes an arrow parameter list.
| Shape | Possible final meaning | What decides |
|---|---|---|
(a, b) | Parenthesized comma expression | When no => follows |
(a, b) => a + b | Arrow parameter list | When the parser later sees => |
({ a } = b) | Object assignment pattern | When an object-looking expression is the assignment target |
({ a: 1 }) | Object literal expression | When parentheses force expression context and no assignment reinterpretation happens |
const samples = [ ["parenthesized expression", "let a = 1, b = 2; (a, b);"], ["arrow parameters", "const fn = (a, b) => a + b;"], ["object assignment pattern", "let a; const b = { a: 2 }; ({ a } = b);"], ["object literal expression", "const value = ({ a: 1 });"],]; for (const [label, source] of samples) { try { new Function(source); console.log(label + ": parses"); } catch (error) { console.log(label + ": " + error.name); }}Object syntax has the same issue. In ({ a } = b), the object-looking shape is reinterpreted as a destructuring assignment pattern. In a normal object literal, { a = 1 } is not a valid property initializer. The spec's cover grammar rules explain why the parser can accept the broad shape first and apply early-error rules only after the final meaning is known.
V8's lazy parsing post discusses exactly this kind of ambiguity for ({ d }: it might become destructuring that references an outer d, or arrow parameters that do not. Modern V8 shares much of the parser and preparser implementation and records scope metadata so later lazy parsing can skip inner functions instead of repeatedly reparsing deeply nested code.
Explore ASTs with real parser output
LIVE ACORNThe fastest way to learn ASTs is to change code and inspect nodes. AST Explorer is the standard web tool. Choose a parser, turn on options, and watch how the tree changes as source changes. In Acorn, Espree, and Babel-style parsers, the options you will reach for first are ecmaVersion, sourceType, locations, and ranges or tokens when a tool needs exact source positions.
const total = prices
.map((price) => price * tax)
.reduce((sum, value) => sum + value, 0);
console.log(total);sourceType: "module"kind: "const"name: "total"optional: falsecomputed: falseoptional: falseoptional: falsecomputed: falseoptional: falsename: "prices"name: "map"
id: nullexpression: truegenerator: falseasync: falsename: "price"operator: "*"name: "price"name: "tax"
name: "reduce"
id: nullexpression: truegenerator: falseasync: falsename: "sum"name: "value"operator: "+"name: "sum"name: "value"
value: 0raw: "0"
optional: falsecomputed: falseoptional: falsename: "console"name: "log"
name: "total"
Acorn parsed the source. Open nodes in the tree and click one to highlight the exact source range that produced it.
Try deleting a closing parenthesis, changing sourceType-sensitive code such as import, or clicking a CallExpression. The playground parses your source with Acorn and highlights character ranges; it never evaluates learner input.
Where frontend tools use ASTs
WALK NODESASTs matter because tools need syntax-aware facts. A linter should flag a real console.log() call without touching the string "console.log". A formatter must understand nesting before it chooses line breaks. A codemod should rename a function call without rewriting comments.
| Tool | How it uses the AST |
|---|---|
| ESLint | Parses source, walks nodes, reports patterns such as unused variables or console.log calls. |
| Prettier | Parses to an AST, then prints a consistent concrete layout from that tree. |
| Babel and SWC | Parse, transform nodes, and generate new JavaScript for a target. |
| Bundlers | Parse imports and exports, build module graphs, and rewrite modules into chunks. |
| Codemods | Find old API shapes in AST form and rewrite them safely across a codebase. |
console.logJavaScriptimport { parse } from "acorn";import { ancestor, full } from "acorn-walk"; const source = `const total = prices.map((price) => price * tax);console.log(total);`;const ast = parse(source, { ecmaVersion: 2024, locations: true }); let identifiers = 0;full(ast, (node) => { if (node.type === "Identifier") identifiers += 1;}); const consoleLogs = [];ancestor(ast, { CallExpression(node) { if ( node.callee.type === "MemberExpression" && node.callee.object.type === "Identifier" && node.callee.object.name === "console" && node.callee.property.type === "Identifier" && node.callee.property.name === "log" ) { consoleLogs.push(node.loc.start.line); } }}); console.log(identifiers);console.log(consoleLogs.join(","));72The helper parses the source, counts every Identifier node with full, and uses ancestor to find a CallExpression whose callee is console.log.
This is the same reason the bundlers lesson used Acorn to find import declarations, and the transpilers lesson used Acorn positions to rewrite ??. Once source is a tree, tools can make precise decisions that plain text matching cannot.
Common misconceptions
- “Grammar and syntax are the same word.” Grammar is the rule system; syntax is the actual source shape being checked against those rules.
- “An AST contains every token.” A parse tree can mirror every grammar detail. An AST keeps a smaller, useful structure.
- “ESTree is what the engine runs.” ESTree is for tools. Engines use their own internal ASTs and compile to bytecode or machine code.
- “An early error happens after earlier lines run.” Early errors prevent evaluation of the entire script, module, or eval code that contains them.
- “Cover grammars mean the parser guesses at runtime.” Cover grammars are a parse-time technique. No JavaScript code executes while the parser reinterprets a shape.
| Pair | Difference | Safe wording |
|---|---|---|
| Syntax vs grammar | Source text vs the rules that accept or reject it | The source has syntax; the parser checks grammar. |
| Parse tree vs AST | Concrete grammar details vs abstract program structure | Tools usually expose ASTs. |
| Early error vs runtime error | Before any evaluation vs while executing a reached statement | Use fixed eval probes to prove which is which. |
| ESTree vs V8 AST | Community tool format vs engine-private representation | Say ESTree-shaped for Acorn output. |
Practice exercises
5 EXERCISESUse the mini parser's precedence rule to predict the root node and operator.
const ast = {
body: [{ expression: { type: "BinaryExpression", operator: "+" } }]
};
console.log(ast.body[0].expression.type);
console.log(ast.body[0].expression.operator);The root is a BinaryExpression for +, so the two logs are BinaryExpression and +.
What ESTree node type should represent a parsed -x expression?
// Extend parsePrimary so a leading '-' becomes:
// { type: "UnaryExpression", operator: "-", argument: parseExpression(30), prefix: true }
// Then remember JavaScript still rejects -2 ** 2 without parentheses.if (peek().value === "-") {
const operator = consume();
const argument = parseExpression(30);
return { type: "UnaryExpression", operator: operator.value, argument, prefix: true };
}A prefix - should produce a UnaryExpression. A production parser must still enforce the exponentiation grammar restriction around **.
Which node type should your visitor inspect first to find real console.log(...) calls?
function isConsoleLogCall(node) {
const callee = node.callee;
return callee.type === "MemberExpression" &&
callee.object.name === "console" &&
callee.property.name === "log";
}
console.log(isConsoleLogCall({
callee: { type: "MemberExpression", object: { name: "console" }, property: { name: "log" } }
}));ancestor(ast, {
CallExpression(node) {
const callee = node.callee;
if (callee.type === "MemberExpression") {
// check console.log here
}
}
});Start from CallExpression nodes because console.log matters only when it is called, not when the property is merely read.
Does the first console.log run before JavaScript reports the duplicate declaration?
console.log("before");
let total = 1;
let total = 2;
console.log("after");Nothing prints. The duplicate let total is an early SyntaxError, so the script never evaluates the first console.log.
For the supported subset, what boolean should the stripped AST comparison print?
const mini = { type: "Program", body: ["same"], sourceType: "script" };
const acorn = { type: "Program", body: ["same"], sourceType: "script" };
console.log(JSON.stringify(mini) === JSON.stringify(acorn));The comparison prints true: after stripping source positions, the mini parser's ESTree-shaped output matches Acorn for config.get(user.name).theme.
Check your understanding
8 QUESTIONSQuestion 1 of 8What does a JavaScript parser consume and produce?
Choose an answer to see the explanation.
Question 2 of 8Why does the mini parser group
a + b * c(d)with*inside+?Choose an answer to see the explanation.
Question 3 of 8What does JavaScript print for right-associative exponentiation?
Read the code, then predictconsole.log(2 ** 3 ** 2);Choose an answer to see the explanation.
Question 4 of 8Which statement about ESTree is accurate?
Choose an answer to see the explanation.
Question 5 of 8What proves an early error prevents earlier lines from running?
Read the code, then predictconst events = []; try { eval('events.push("before"); let x; let x;'); } catch (error) { events.push(error.name); } console.log(events.join(" -> "));Choose an answer to see the explanation.
Question 6 of 8Why can
(a, b)be a cover grammar shape?Choose an answer to see the explanation.
Question 7 of 8Which parser option tells Acorn whether top-level
importis allowed?Choose an answer to see the explanation.
Question 8 of 8Which tool workflow is AST-based?
Choose an answer to see the explanation.
Key takeaways
- The parser consumes tokens, checks grammar, reports syntax and early errors, and builds a tree.
- Recursive descent uses functions for grammar rules; Pratt parsing adds binding power for expressions.
- ESTree is a community AST format for tools, not V8's internal AST.
- Early errors prevent any evaluation of the script, module, or eval code that contains them.
- Cover grammars let the parser accept broad shapes and reinterpret them when later context arrives.
- ASTs power linting, formatting, transforms, codemods, bundling, and clearer syntax error messages.
Remember the one-liner.
Parsing is where JavaScript stops being a token stream and becomes a checked tree that engines and tools can reason about.
Up next: Lazy parsing & preparsing, where engines decide how much of that work to do during startup and how much to defer.