Lexical grammar & ASI
Read how ECMAScript treats whitespace, names, literals, comments, and line breaks, then predict when automatic semicolon insertion changes code.
- 01Read source boundariesDistinguish white space, line terminators, comments, tokens, and lexical goals in the specification.
- 02Recognize legal spellingsExplain Unicode identifiers, reserved words, literal separators, hashbangs, and trailing commas.
- 03Predict ASIApply its three rules and spot restricted productions and continuation hazards before changing app code.
Source text and lexical grammar
The previous lesson, A tour of the abstract operations, follows named rules after source has a meaning. Here we start one step earlier. Lexical grammar describes how source characters form input elements: tokens, white space, comments, and line terminators. Automatic semicolon insertion (ASI) is the narrow set of rules that sometimes supplies a missing semicolon while that source is parsed.
const price = 3;console.log(price + 2);Line 1 stores 3 in price. Line 2 adds 2 and prints 5. The spaces separate names and punctuation; the semicolons are written explicitly. We will keep this kind of small result visible while changing one piece of spelling at a time.
A token is one lexical unit, such as price, 3, or +. The spec scans source left to right and generally takes the longest possible input element. The surrounding syntax decides which lexical goal is used where the same character could begin different tokens.
This is the standard's view of source, not a tour of a particular engine's tokenizer. For implementation details, visit Scanner & tokens. Read the live ECMAScript lexical grammar alongside this lesson; the next lesson follows the syntactic grammar that consumes its tokens.
Format-control characters and white space
WhiteSpace separates tokens but does not itself finish a statement. A space, tab, and the U+FEFF byte order mark (BOM, also called zero-width no-break space) are in this group. A BOM can occur within source, not just at the start of a file. Outside a literal or comment the spec treats it as white space.
constprice = 3;console.log(price);Line 1 has an invisible U+FEFF between const and price. It separates the two tokens as a space would. Line 2 prints 3. When copying source from another file, an invisible character may be worth inspecting even if the code still parses. Inside a string the same code point is data, not a separator.
const price = 3; /* notefor the reader */ console.log(price);Line 1 stores the price and starts a comment. Line 2 closes the comment and prints 3. A multi-line comment containing a line terminator counts as a line terminator to the surrounding syntax; a comment without one behaves like white space. The characters of this comment never become program tokens.
| Input | Role | Example |
|---|---|---|
| Space, tab, BOM | WhiteSpace: separates tokens, does not force ASI. | A space between const and price |
| LF, CR, LS, PS | LineTerminator: separates tokens and can affect ASI. | The line break after return |
| Single-line comment | Discarded; its ending line terminator is separate. | // note followed by a new line |
| Multi-line comment | Discarded; counts as a line terminator when it contains one. | /* note followed by a new line and */ |
Format-control characters other than BOM are not generally free spacing in code. They are allowed within comments and literals, and two particular ones are allowed inside identifiers as continuation characters. We will meet those joiners after looking at line endings.
Line terminators and comments
A LineTerminator is a source line break recognized by the grammar: line feed (LF), carriage return (CR), line separator (LS, U+2028), or paragraph separator (PS, U+2029). CR followed by LF forms one line terminator sequence. Unlike an ordinary space, a terminator can affect ASI and a production marked [no LineTerminator here].
const price = 3;console.log(price + 2);Line 1 stores 3. Line 2 starts a print call, and line 3 adds 2 before closing it. The output is 5. A line break inside this expression is legal; the parser does not insert a semicolon merely because the text looks like a new line to us.
const note = 'a b';console.log(note.length);Line 1 stores a string with a real tab between a and b. Line 2 prints 3, the string's length. White space inside a literal is part of its value; it does not separate JavaScript tokens there.
const line = 'tea
milk';console.log(line.length, line.includes('
'));Line 1 includes a literal U+2028 between tea and milk. Line 2 prints 8 true: the code point belongs to the string. Since ES2019, LS and PS can appear unescaped in a quoted string; LF and CR still cannot, except as part of a backslash line continuation that contributes no character.
A // comment ends before its line terminator, leaving that terminator available to the grammar. A /* ... */ comment that contains a line break counts as a line terminator. This matters when reading a return statement later: a comment does not erase the break inside it.
Identifiers and Unicode escapes
An IdentifierName is a legal sequence of name characters before the syntactic context decides whether it can act as an identifier. Unicode ID_Start characters, $, and _ can begin one. After the first character, ID_Continue characters can follow. The actual sequence of code points matters, not how the name looks on screen.
const café = 3;console.log(café);Line 1 declares café with value 3. Line 2 reads exactly that spelling and prints 3. A different Unicode normalization is not made equal merely because a human sees the same letters.
const c\u0061rt = 3;console.log(cart);Line 1 spells the letter a with \u0061, so the binding's name is cart. Line 2 reads it without an escape and prints 3. The backslash does not add a character to the name. An escape is not a loophole: it cannot put a digit at the start of a name or disguise a reserved word.
const tea\u200Ccup = 3;console.log(tea\u200Ccup);Line 1 puts U+200C (zero-width non-joiner, ZWNJ) between tea and cup; line 2 reads the same name and prints 3. U+200D (zero-width joiner, ZWJ) is allowed in the same continuation position. Neither is an identifier start, and neither is ordinary spacing between tokens.
Use plain names for app code whenever possible. The point of this rule is to understand why some pasted names are valid but visually surprising. The next section separates a name's legal characters from its legal role in a program.
Reserved words and contextual keywords
A reserved word cannot be used as an ordinary identifier in the relevant context. For example, return is reserved, but an object property after a dot can have the IdentifierName default. A contextual keyword has a special use only in some grammar positions. async can be a binding name.
const async = "tea";const cart = { default: async };console.log(cart.default);Line 1 binds async to tea. Line 2 stores that value in the property named default. Line 3 prints tea. Here neither spelling starts an async function or attempts to declare a variable named default.
const \u0065lse = 3;console.log("unreachable");Line 1 tries to spell else by escaping its first letter. It is a SyntaxError before line 2 can log. In contrast, await and yield are context dependent; for example, await cannot be an identifier in a module or inside an async function. Strict mode also restricts some additional names.
| Name | Meaning | Example |
|---|---|---|
| IdentifierName | Spelling accepted by lexical grammar | default after a dot can be a property name. |
| Identifier | IdentifierName allowed as a binding or reference here | const price = 3 is valid; const default = 3 is not. |
| Contextual keyword | A spelling that has special syntax in some places | async is also a legal variable name. |
This distinction is useful when reading a spec production: the lexical grammar can recognize an IdentifierName, while the syntactic grammar or an early error still rejects its use as a binding. Next we turn to literals, whose spelling also changes what can be written without changing the underlying value.
Literals and numeric separators
A literal is source spelling for a value: a number such as 3, a quoted string, or a regular expression such as /tea/. A numeric separator is an underscore between digits in a number. It makes a large number easier to read but contributes no value.
const price = 1_000;console.log(price + 2);Line 1 stores one thousand in price. Line 2 adds 2 and prints 1002. The underscore is inside the numeric literal, not a separate JavaScript operator.
const cart = 0b1010;const total = 10n;console.log(cart, typeof total);Line 1 reads binary 0b1010 as 10. Line 2 creates the BigInt 10n. Line 3 prints 10 bigint. String and template literals have their own escape rules; numeric separators belong to numeric spelling, not string parsing or Number("1_000").
const price = 1__000;console.log(price);Line 1 is invalid; line 2 cannot run. Separators must separate digits, not sit together, at the beginning or end, beside a decimal point or exponent marker, or immediately after a base prefix such as 0x. Our tests check those spellings as syntax errors instead of executing them in the page.
The three automatic semicolon insertion rules
ASI does not insert a semicolon after every newline. It inserts one in three defined cases while parsing. First, when an offending token cannot continue the grammar, it may insert before that token if a line terminator separates it from the previous token, if it is }, or after a do-while's closing ).
const price = 3+ 2console.log(price)Line 1 starts a declaration with 3; line 2 continues that value with + 2. Line 3 begins a new statement that the parser cannot attach to the declaration, so the code prints 5. It would be wrong to say a semicolon was inserted before the plus.
Replay of instrumented example code; this is not an engine debugger.
script
+ 2console.log(price)The replay records the lesson's addition, not a parser trace. Step through the saved value and its print result. Second, if the token stream ends and the whole program otherwise cannot parse, ASI inserts at end of input. Third, if a token follows a [no LineTerminator here] position across a line break, ASI inserts before that restricted token. The next section gives that rule a visible result.
console.log(3)The only line prints 3 without a written final semicolon. A closing brace can also allow insertion before it, and a do-while has its special after-) case. ASI never supplies an empty statement or either of the two semicolons required inside a for header.
| Rule | When insertion is considered | Source example |
|---|---|---|
| Offending token | Insert before a token that cannot continue the parse, if separated by a line terminator, if the token is }, or after ) closing a do-while. | const price = 3 then console.log(price) |
| End of input | Insert at the end when the program otherwise cannot parse as a whole. | A final console.log(3) |
| Restricted token | Insert before a token separated by a line terminator where [no LineTerminator here] applies. | return then price |
Restricted productions: keep these together
A restricted production has a spot marked [no LineTerminator here]. A line break after return, throw, break, continue, or yield changes how the following token is read. The exact outcomes differ: a bare return is legal inside a function, but a bare throw followed by its value is a SyntaxError.
function receipt() { const price = 3; return void price;}console.log(receipt());Line 1 starts receipt; line 2 stores 3. Line 3 returns with no value because the line terminator forbids line 4 from becoming the argument. Line 4 is a separate, unreachable expression. Line 6 prints undefined, not 3.
Replay of instrumented example code; this is not an engine debugger.
script
function receipt() { const price = 3; return void price;}A teacher writes an instruction on one notebook line. A number on the next line is a new entry, not part of that instruction.
- In real life: A line ends a written instruction
- In JavaScript: A restricted line ends a return
- In real life: The next line begins separately
- In JavaScript: The following expression is separate
- In real life: A mark belongs on its line
- In JavaScript: Return values belong on the return line
Where the analogy stops: Most JavaScript line breaks do not end statements; this analogy applies only to restricted positions.
let score = 3;score++score;console.log(score);Line 1 sets score to 3. Line 2 is a standalone read; line 3 is a prefix ++score, not a postfix update on line 2. Line 4 prints 4. The same rule applies to postfix --. A line break before an arrow's => instead causes a SyntaxError.
const show = (price)=> price;console.log(show(3));Line 1 closes the parameter list. Line 2 begins with an arrow that cannot attach across the line break, so no log runs. Similarly, async must share a line with the following function, identifier, or ( when it is part of an async definition or arrow. Otherwise it is a separate identifier in contexts that allow it.
function receipt() { throw new Error("missing");}Line 2 cannot be a bare throw; line 3 cannot repair it. break and continue also require any label on the same line, while yield can become a bare yield with a later expression separate. Keep the operand or label on the same line whenever the grammar requires it.
Lines that continue an expression
Sometimes the next token can continue the previous expression. Then no ASI rescue happens. A new line beginning with ( can call the previous value. If that value is a number, the result is a TypeError, not a SyntaxError. Try the fixed two-source playground before changing a real file.
const price = 3(function addFee() { return 2; })()console.log(price);Line 1 leaves 3 as the previous expression. Line 2 supplies what parses as its argument list; the program tries to call 3. Line 3 never prints. The result is a TypeError saying the value is not a function. Adding a semicolon after line 1 makes these independent statements.
const price = 3(function addFee() { return 2; })()console.log(price);TypeError: 3 is not a functionThe parenthesis joins the previous expression, attempting to call the number 3. That throws a TypeError before the log.
The control changes exactly one character in the displayed source. Reset restores the version with no semicolon. The panel labels the expected result for these fixed inputs; Edit & run executes the code itself.
const price = 3;(function addFee() { return 2; })();console.log(price);Line 1 ends the declaration explicitly. Line 2 calls its own function, which returns 2 without printing it. Line 3 prints 3. The change fixes a parse boundary, not the value of price.
const price = 6 / 2;const hasTea = /tea/.test("tea");console.log(price, hasTea);Line 1 uses / as division and gets 3. Line 2 uses a RegExp literal and gets true from test. Line 3 prints 3 true. The spec calls the lexical goals InputElementDiv and InputElementRegExp: syntactic context determines which slash spelling can start a token. ASI does not simply choose a RegExp because a slash begins a new line.
| Start | Possible continuation | When a new statement was intended |
|---|---|---|
( | A call of the prior expression | End the prior statement with ; before an intended grouping. |
[ | Property access on the prior expression | Use ; before an intended array literal. |
| backtick | A tagged template with the prior expression | Use ; before an intended template. |
+ or - | A binary operator joining the lines | Use ; when the next line starts a new expression. |
/ | Can become division rather than a RegExp literal | Use ; before an intended regular expression. |
A leading [ can mean property access; a backtick can form a tagged template; + and - can be binary operators. A leading slash may be division, depending on the lexical goal. The safe habit is to end the previous expression with ; before a line intended to start with any of these characters, or explicitly prefix that line with a semicolon.
Hashbang comments and trailing commas
A hashbang comment starts with #! at the very beginning of a Script or Module. It is discarded like a comment. Its special lexical goal only applies at that starting position; later #! is not another hashbang comment. Some command-line hosts use it to pick an interpreter.
#! sample scriptconsole.log("tea");Line 1 is the hashbang and produces no output. Line 2 prints tea. Put nothing, not even another comment, ahead of it in the source text. The hashbang is a lexical exception, not a new JavaScript operator.
const cart = ["tea", "milk",];const order = { price: 3, };console.log(cart.length, order.price);Line 1 gives the array two items with a trailing comma. Line 2 gives the object a price field with a trailing comma. Line 3 prints 2 3. These commas are allowed by the surrounding syntactic productions, not inserted by ASI; a comma is not a semicolon.
const cart = ["tea",];const withHole = ["tea",,];console.log(cart.length, withHole.length);Line 1 leaves a single element despite its comma. Line 2 has two commas after the value, which creates an elision (an empty slot) before the trailing comma. Line 3 prints 1 2. Do not infer array length from commas alone without noticing whether an element is missing between them.
Trailing commas also appear in supported parameter and argument lists, but JSON is a different grammar and rejects them. Use a JavaScript parser for JavaScript source; do not assume JSON has all its punctuation rules.
Review a real app change
Imagine a checkout page builds an item list and then adds a second item on a new line. During a refactor, a line beginning with [ can be read as indexing the previous expression. Put a clear boundary before it, then check the value in the console.
const total = 3;const show = (value) => console.log(value);show(total);Line 1 stores 3. Line 2 defines show to print its argument. Line 3 calls show with total, so it prints 3. Keeping an arrow's parameter and => together prevents an accidental restricted-line error during this refactor.
const items = ["tea"];const count = items.length;;["milk"].forEach((item) => items.push(item));console.log(count, items.length);Line 1 creates a list with one item. Line 2 records length 1. Line 3 explicitly begins a new statement with a semicolon and pushes milk. Line 4 prints 1 2. The leading semicolon is useful when the previous line belongs to generated or concatenated code you cannot edit.
- A space between
constandprice - U+FEFF outside a string
- LF after
return - U+2029 between statements
- The name
price - The literal
1_000
Sort each example as white space, line terminator, or token; then read why the choice matters.
In code review, first ask whether a line is part of the previous expression. Then check for a restricted line break, and finally check any special spelling such as a separator or escape. That is more reliable than adding semicolons by guessing where the parser might put them.
Common misconceptions
A printed result is better evidence than a rule of thumb about newlines. Read where the next token can attach, and remember that errors are observable too. A SyntaxError stops parsing before the first log; a TypeError in the leading-parenthesis example happens while running the parsed program.
const price = 3+ 2console.log(price)Line 1 starts the price declaration, line 2 continues it with addition, and line 3 prints 5. There is no inserted semicolon before the plus. Compare that with return and a newline: its restricted position changes the result even when an expression follows.
- "ASI means one semicolon per newline." Only the three specified rules insert one.
- "Spaces and line breaks are the same." Only line terminators trigger restricted positions.
- "An escape makes a reserved word safe." The decoded identifier still gets checked.
- "An underscore works anywhere in a number." It must occur between suitable digits.
- "A slash on a new line must be a regex." The syntax chooses a lexical goal first.
- "Trailing commas are ASI." They belong to the grammar for arrays, objects, and other lists.
A special case matters here: a line terminator inside a multi-line comment still counts. If that comment sits after return, the next expression cannot become the return value. Comments are not a way to hide source boundaries from the syntactic grammar.
Practice exercises
Predict before running. For each snippet, name the exact character that matters: an escape, an underscore, a line terminator, or an explicit semicolon. Type the requested answer, then compare the worked explanation.
A receipt function should return the price, but a formatting change moved it down. Predict what the starter prints. Then inspect the worked fix.
function receipt() {
const price = 3;
return
void price;
}
console.log(receipt());function receipt() { return 3; }
console.log(receipt());The starter prints undefined. To return 3, put the value on the same line as return, as the worked code does.
A scoreboard starts at three. Read the two update lines as the grammar does, then type the final score. Does a postfix operator cross the line?
let score = 3;
score
++score;
console.log(score);let score = 3;
score
++score;
console.log(score);The starter prints 4. The newline separates the read of score from the prefix ++score.
You paste a small cart snippet from another file. Predict its output after decoding the identifier escape. Check the real spelling before renaming anything.
const c\u0061rt = 3;
console.log(cart);const c\u0061rt = 3;
console.log(cart);It prints 3. c\u0061rt and cart name the same binding.
The checkout price is written with a digit separator. Type the printed number. Explain why the underscore has no numeric value.
const price = 1_000;
console.log(price + 2);const price = 1_000;
console.log(price + 2);The output is 1002. An underscore within the valid literal changes readability, not value.
A checkout calculation unexpectedly calls a number. Name the punctuation that makes the intended two statements separate. The starter is intentionally broken.
const price = 3
(function addFee() { return 2; })()
console.log(price);const price = 3;
(function addFee() { return 2; })();
console.log(price);Put a semicolon after 3. Now the next line is its own function call and the final log prints 3.
Your site appends an item in a separate statement. A previous line may later lose its ending semicolon. What would you add immediately before a line starting with an array?
const items = ["tea"];
const count = items.length;
;["milk"].forEach((item) => items.push(item));
console.log(count, items.length);const items = ["tea"];
const count = items.length;
;["milk"].forEach((item) => items.push(item));
console.log(count, items.length);Add an explicit semicolon before ["milk"]. This keeps the list update independent of the preceding expression; the program prints 1 2.
Check your understanding
Read each code question as source, not as a visual list of lines. Check whether the next token continues an expression, and whether the grammar marks a restricted break. When a choice names an error, decide whether parsing failed or execution reached a bad call.
Question 1 of 7What does the receipt function print?
Read the code, then predictfunction receipt() { return 3; } console.log(receipt());Choose an answer to see the explanation.
Question 2 of 7What does an underscore do in
1_000?Read the code, then predictconsole.log(1_000 + 2);Choose an answer to see the explanation.
Question 3 of 7Which whitespace can affect automatic semicolon insertion?
Choose an answer to see the explanation.
Question 4 of 7What does this contextual keyword example print?
Read the code, then predictconst async = "tea"; console.log(async);Choose an answer to see the explanation.
Question 5 of 7What does the leading plus do here?
Read the code, then predictconst price = 3 + 2 console.log(price)Choose an answer to see the explanation.
Question 6 of 7Can an escape turn an invalid name into a valid identifier?
Choose an answer to see the explanation.
Question 7 of 7Which describes a hashbang comment?
Choose an answer to see the explanation.
The quiz uses small programs you can run, plus rule questions grounded in the spec. For a wrong choice, read its explanation and revisit the associated code block before retrying.
Key takeaways
Source code has structure before it runs. White space, line terminators, comments, and tokens are different input elements. ASI uses the grammar and its restrictions; formatting alone is not enough to predict it.
- BOM is white space outside literals; LF, CR, LS, and PS are line terminators.
- Unicode escapes contribute identifier code points but cannot hide an invalid name or reserved word.
- Numeric separators sit between digits; hashbangs belong only at the start of source.
- ASI has three rules and never inserts every possible semicolon or repairs a for header.
- Keep return values, throw values, labels, postfix operators, arrows, and async heads together where required.
- Before lines beginning with (, [, backtick, +, -, or /, write the intended statement boundary explicitly.
Remember the one-liner.
A newline changes syntax only when a grammar rule makes it matter.
Coming next: Syntax, early errors & static semantics. That lesson explains the grammar shapes and checks that follow token recognition.