Benchmarking without fooling yourself
Learn to warm up JavaScript, prevent discarded work, compare fixed sample distributions, and choose realistic browser benchmark suites.
- 01Separate phasesExplain why early interpreter work and later optimized work need separate warm-up and measurement phases.
- 02Keep work observableWrite benchmark code whose result is used, so an optimizer cannot remove the work you meant to compare.
- 03Report uncertaintyUse many samples, median, spread, and overlapping ranges before saying one implementation is faster.
What a benchmark can and cannot say
A microbenchmark is a very small program used to compare one narrowly defined piece of work. It can help answer a focused question. It cannot, by itself, predict whether a whole page feels fast.
const start = performance.now();let sum = 0;for (let i = 0; i < 100000; i++) sum += i;console.log(sum);Line 1 reads a clock. Line 2 creates sum. Line 3 adds the numbers from 0 through 99,999. Line 4 prints the fixed result 4999950000. The lesson does not print the elapsed time because it changes on every run, device, and browser state.
That difference matters. The calculation has a stable behavior-level result, which tests can prove. A duration is evidence about one environment at one moment. Treat it as a sample to investigate, not a fact to memorize.
A useful benchmark note begins with a question a teammate can disagree with. “Is formatting an already-loaded cart total slower after this change?” is a question. “Which JavaScript trick is fastest?” is too broad, because it does not name data size, browser, phase, or user task.
Small code can still matter, but only in context. A cart formatter called once after a network request has different importance from one called while a user types on every keystroke. The call count, input shape, and place on the critical path decide whether a tiny difference deserves attention.
Benchmark code should state what it measures, what it does not measure, which phase it measured, and how uncertain the comparison is.
Start with a performance measurement question from a real page. Then use a microbenchmark only when you can isolate one cause without confusing it with layout, network, input, or framework work.
Warm-up and JIT tiers
JavaScript engines often begin a function in an interpreter. When they see a stable pattern many times, they may compile it into a faster tier. This is tiered compilation: the engine changes how it runs the same JavaScript while keeping its visible behavior.
function addCartPrices() { let total = 0; for (let i = 0; i < 3; i += 1) total += i; return total;} for (let run = 0; run < 2; run += 1) addCartPrices();console.log(addCartPrices());Lines 1 through 5 define a function that always returns 3. Line 7 calls it twice as warm-up. Line 8 prints 3 from another call. The output proves the function result, not a speed claim.
The first cup of tea can be slow because you find the tea and heat the kettle. Later cups can be quicker. Timing only the first cup or only the hundredth cup answers different questions.
- In real life: The first cup finds tea and heats water
- In JavaScript: Early calls can include interpreter and setup work
- In real life: Later cups follow a familiar routine
- In JavaScript: Later calls may reach a faster JIT tier
- In real life: A notebook labels which cups were counted
- In JavaScript: A report labels warm-up and measured phases
Where the analogy stops: An engine does not literally make tea, and it may change tiers differently across browsers.
Do not call later measurements “the JavaScript speed.” They are measurements after a chosen warm-up policy. Link this model with tiered compilation to learn why engines choose and leave optimized tiers.
A cold-start benchmark can be valid when cold start is the product question. A steady-state benchmark can be valid when a long-running editor or animation is the question. The error is hiding which question the benchmark answered.
Tier changes are not a contract from JavaScript. A new input shape, a changed object layout, a browser update, or a different engine can produce a different path. That is why a benchmark report records the browser and its version instead of treating one local result as a language rule.
A small honest harness
A benchmark harness is the small program that runs work, collects samples, and reports them. The harness can accidentally change the result, so keep it simple and make its phases visible.
function measure(work, clock, warmupRuns, measuredRuns) { for (let run = 0; run < warmupRuns; run += 1) work(); const samples = []; for (let run = 0; run < measuredRuns; run += 1) { const start = clock(); work(); samples.push(clock() - start); } return samples;} const clock = makeFakeClock([10, 13, 20, 24]);console.log(measure(() => 6, clock, 1, 2).join(","));Line 2 runs warm-up calls before creating any samples. Line 3 creates the measured list. Lines 5 through 8 read the injected clock around one work call and store a difference. The shown fake clock makes line 13 print 3,4.
The fake clock is not a shortcut for a real benchmark. It lets this lesson test the harness logic exactly. A real harness uses a real high-resolution clock, repeats more samples, and reports its environment.
Keep setup outside the measured region when setup is not part of the question. Parsing test data, creating a large array, opening a database, and rendering a button can all be important work. Include them only when the user-facing task includes them, then say so in the benchmark name.
This is a replay of instrumented lesson code. It shows which phase is warm-up and which values become measurements; it is not an engine debugger.
script
for (let run = 0; run < warmupRuns; run += 1) work(); const samples = []; for (let run = 0; run < measuredRuns; run += 1) { const start = clock(); work(); samples.push(clock() - start); } return samples;} const clock = makeFakeClock([10, 13, 20, 24]);console.log(measure(() => 6, clock, 1, 2).join(","));Notice what the replay does not say: it never claims your browser took 3 ms or 4 ms. It demonstrates phase separation. That is the behavior-level claim we can prove without a timing race.
Dead-code elimination
Dead-code elimination is an optimization that can remove work whose result cannot affect what anyone can observe. That is normally helpful. It is dangerous when a benchmark accidentally measures a loop that no longer has meaningful work.
let total = 0;for (let i = 0; i < 4; i += 1) { Math.sqrt(i); // discarded work}console.log(total);Line 1 creates total. Line 3 calculates a square root and immediately throws its value away. Line 5 prints 0, because no line changes total. A modern optimizer may be able to treat discarded calculations very differently from useful work.
let total = 0;for (let i = 0; i < 4; i += 1) { total += Math.sqrt(i);}console.log(total.toFixed(3));Line 3 now adds each square root into total. Line 5 prints 4.146. The result is observable, so changing or removing the computation would change the program. This proves useful behavior, not a timing.
Sometimes a benchmark stores a result in a sink instead of printing it. The important rule is the same: use the result in a way the runtime must preserve. Do not rely on a loop body merely looking expensive in source code.
Do not overcorrect by adding logging inside every measured iteration. Console output changes the work being measured. A good sink is quiet during the loop and observed after the loop, such as one final total, a returned checksum, or a value consumed by the calling code.
Many samples, not one lap
One lap time can be luck. Another process may run, the CPU may change state, or the engine may perform runtime work. A benchmark should collect many samples and show their shape instead of choosing the most flattering number.
const samples = [10, 11, 10, 12, 50];console.log(mean(samples));console.log(median(samples));console.log(spread(samples).toFixed(2));Line 1 stores five fixed laps. Line 2 prints the mean 18.6. Line 3 prints the median 11. The one slow 50 moves the mean far more than the middle value. A spread tells you how far samples sit from their average.
A report should not choose either number as a mascot. The mean answers “what was the average across these samples?” The median answers “what was the middle sorted sample?” A reader needs both questions, plus the slow edge, to see whether one long pause changed the story.
| Number | What it helps you see | What it can hide |
|---|---|---|
| Warm-up | Runs before recording samples | Interpreter work or a tier transition can dominate it |
| Measured run | A run included in the report | It still varies because the machine and runtime do other work |
| Mean | All samples added and divided by count | One very slow outlier can move it strongly |
| Median | Middle sorted sample | It hides the size and frequency of slow outliers |
| Suite | A realistic group of workloads | It cannot prove one tiny code expression always wins |
Time many laps and look at the middle and the spread. One lap can be lucky or unlucky. Keeping the whole group stops one surprising lap from quietly becoming the story.
- In real life: One lap can have a red light
- In JavaScript: One sample can be an outlier
- In real life: The middle lap shows a typical trip
- In JavaScript: The median resists one extreme sample
- In real life: The slowest and fastest laps stay in the notebook
- In JavaScript: Range and spread show variation
Where the analogy stops: Real benchmark samples are not car laps; statistics describe evidence but do not explain its cause.
This replay uses fixed sample arrays. It demonstrates the statistics a benchmark report should include without pretending to measure this device.
script
const withOutlier = [10, 11, 10, 12, 50];const before = summarizeSamples(withoutOutlier);const after = summarizeSamples(withOutlier);console.log(before.mean, before.median);console.log(after.mean, after.median);Median is not magic. If many laps are slow, the median moves too. Show the count, minimum, maximum, median, mean, and spread so another developer can see both the typical case and the unstable edges.
Also avoid deleting inconvenient samples without a rule written in advance. A pause might be noise, or it might be the exact user-visible problem you need to explain. Keep the raw list when it is small enough, state exclusions when they exist, and investigate surprising values rather than quietly hiding them.
Read a fixed sample report
This playground changes one input: whether the fixed 50 sample is present. It does not read your clock. That makes the effect of an outlier inspectable and resettable.
const samples = [10, 11, 10, 12, 50];console.log(mean(samples));console.log(median(samples));console.log(spread(samples).toFixed(2));10, 11, 10, 12, 50The observed values in this teaching model.
18.6 / 11The mean moves with the outlier; the median is the middle sorted value.
10 / 50The full range prevents a calm middle number from hiding a slow lap.
15.72Standard deviation is a simple measure of how far values sit from their mean.
The report uses fixed sample arrays. With this choice, the comparison verdict is not sure. Reset restores the outlier.
When two reported ranges overlap, this teaching model says not sure. That is not a failure. It is a useful result: collect more representative samples, reduce noise, or use a profiler on the real interaction.
Range overlap is intentionally a simple rule here, not a statistical proof. Real performance work may use confidence intervals, more repetitions, and carefully controlled environments. The practical habit is still simple: do not claim a winner when your own samples do not clearly separate.
Speedometer and JetStream
Microbenchmarks are narrow by design. Browser benchmark suites answer broader, still limited questions with collections of workloads. They are useful evidence when you understand what each suite includes.
const suite = { speedometer: "web app responsiveness", jetStream: "JavaScript and WebAssembly",};console.log(suite.speedometer);console.log(suite.jetStream);Lines 1 through 4 create two labels. Line 5 prints web app responsiveness. Line 6 prints JavaScript and WebAssembly. Those are fixed labels, not benchmark scores.
Speedometer 3 measures web application responsiveness by timing simulated user interactions across workloads. It is developed under a multistakeholder model by the three widely distributed browser engine projects. Its documentation warns against using it as a framework ranking.
JetStream is a JavaScript and WebAssembly suite focused on advanced web applications. It rewards quick startup, execution, and smooth running. Both suites are evidence about their workloads, versions, and test environment, not universal laws.
Use a suite when you want a broad browser signal, such as a regression check across browser versions. Use a profile when you want to know why your own interaction is slow. Use a tiny benchmark only when a profile has already narrowed the question to one small choice.
Benchmark a real web change
Suppose a checkout page feels slow after the user changes quantity. Start by recording the real interaction with the CPU profiler. Find out whether time is in your JavaScript, rendering, data work, or a network wait before writing a benchmark.
Then state one narrow question, such as whether two ways to format an already-loaded cart total differ after warm-up. Keep input data fixed, use the returned value, run enough samples, and keep the page-level profiler result beside the microbenchmark.
Write down the browser version, operating system, device class, input size, warm-up count, measured count, and whether the tab was quiet. Those details do not make the result universal. They make it possible for a teammate to understand what your result actually describes.
After the code change, repeat the user flow. A tiny benchmark can improve while the page gets worse because allocation, rendering, or a different call path changed. The product interaction is the final check.
Keep the benchmark code in review with the change it informed. Future readers need the question, inputs, and limits as much as they need the result.
- Write the product question before writing the loop.
- Warm up and label whether you measure cold start or steady state.
- Keep the calculated result observable.
- Report distributions and environment details, not one winning number.
- Check the real interaction again before merging an optimization.
- Run the same cart-total function before storing samples.
- Add square-root results to
totaland print it. - Read the fake clock before and after one work call.
- Store several post-warm-up samples.
- State the median alongside the mean.
- Say
not surewhen the two ranges overlap.
Sort each step into the phase where it belongs. The explanation tells you why the step protects an honest conclusion.
Common benchmark traps
- “The fastest single run wins.” One run is a sample, not a conclusion.
- “Warm-up is cheating.” It is needed when steady-state behavior is the question; label it honestly.
- “The loop obviously does work.” A discarded result may not represent work an optimizer must keep.
- “The mean settles every argument.” Outliers and overlap require more context.
- “A suite score is my page score.” Suites model workloads; profile the page your users use.
| Idea | Meaning | Important limit |
|---|---|---|
| The first number is the answer | Early calls often include cold setup and interpreter execution | Warm up, then label the phase you measured |
| A loop always measures its body | Unused results may be removable work | Use the result through a printed total or sink |
| A lower mean proves faster | The distributions may overlap or have different spread | Report median, range, and a cautious verdict |
| A microbenchmark predicts app speed | An app includes DOM, layout, network, input, and framework work | Profile the real interaction and use suites for broad comparisons |
Practice exercises
Read the short program and type its console output.
let total = 0;
for (let i = 0; i < 4; i += 1) total += i;
console.log(total);It prints 6: 0 + 1 + 2 + 3. A fixed output is suitable for a lesson test; elapsed time is not.
What is the name of the phase that runs before a harness stores samples?
The warm-up phase runs before measurements. State whether your report measures cold start or post-warm-up work.
Your loop calculates a value for each cart item. What should happen to that calculated result?
Use the result: accumulate and print it, return it to a caller, or store it in a sink. The optimizer must preserve observable behavior.
What does the fixed median program print?
const samples = [10, 11, 10, 12, 50];
console.log([...samples].sort((a, b) => a - b)[2]);It prints 11. The 50 ms outlier remains visible in the maximum, but it does not become the middle sample.
What should you say when two small sample ranges overlap?
console.log(rangesOverlap([10, 12], [11, 13]) ? "not sure" : "faster");The output is not sure. More evidence or a real profile is needed before claiming one choice is faster.
A quantity change feels slow in a real checkout page. What should you inspect before replacing its cart-total code?
Profile the real interaction first. Only then isolate a narrow question, benchmark it honestly, and check the real checkout flow again.
Check your understanding
Choose answers that distinguish stable program behavior from measurements that need context.
Question 1 of 7Why should an honest microbenchmark warm up before it stores measurements?
Choose an answer to see the explanation.
Question 2 of 7What does this fixed-output loop print?
Read the code, then predictlet total = 0; for (let i = 0; i < 4; i += 1) total += i; console.log(total);Choose an answer to see the explanation.
Question 3 of 7Why is
total += Math.sqrt(i)safer than discardingMath.sqrt(i)in a benchmark?Choose an answer to see the explanation.
Question 4 of 7The samples are 10, 11, 10, 12, and 50. Which number best shows the typical lap here?
Choose an answer to see the explanation.
Question 5 of 7Two reports have ranges [10, 12] and [11, 13]. What should this teaching model say?
Choose an answer to see the explanation.
Question 6 of 7What does Speedometer measure?
Choose an answer to see the explanation.
Question 7 of 7What does JetStream include?
Choose an answer to see the explanation.
Key takeaways
- Measure a named question, not “JavaScript speed.”
- Warm up when the question is steady state, and label that phase.
- Use benchmark results so discarded work cannot define the test.
- Show many samples, median, spread, and uncertainty.
- Use Speedometer and JetStream as workload-scoped evidence, not universal rankings.
- Profile real user interactions before and after optimization.
Remember the one-liner.
A benchmark is honest when its work, phase, samples, and limits are visible.
Coming next: The HTML event loop processing model, where browser tasks, microtasks, and rendering steps decide what runs next.