Measuring performance
Learn to profile JavaScript bottlenecks with Performance API marks, DevTools flame charts, RUM data, and safer benchmarks.
- 01Measure the right speedUse percentiles, User Timing marks, navigation timing, and observers before choosing an optimization.
- 02Read profiles instead of guessingNavigate Chrome DevTools recordings, main-thread tasks, flame charts, Bottom-up, Call tree, Event log, and Insights.
- 03Compare lab, field, and benchmarksSeparate Lighthouse-style lab runs, CrUX and RUM field data, and careful microbenchmarks with warm-up and medians.
Measure before you make it fast
Professional JavaScript performance work starts with evidence. A slow feeling might come from network latency, parser work, a long task, layout, a third-party script, or a benchmark that never matched production. Guessing wastes time and can make the page feel worse.
Measuring performance means collecting trustworthy timing evidence from the browser, DevTools, real users, and carefully designed benchmarks before deciding what to optimize.
Keep the scope clear. This lesson teaches measurement: User Timing, DevTools traces, flame charts, RUM, and benchmark traps. The next lesson, Keeping the main thread responsive, goes deeper on long tasks, INP, and yielding. Later lessons cover debounce and throttle, rendering performance, loading performance, and memory leaks.
A good doctor does not prescribe random medicine because someone says “I feel tired.” They check vitals, choose a test, read the result, and decide what treatment fits. Performance work is the same: measure first, then optimize the bottleneck you actually found.
- In real life: A symptom starts the investigation
- In JavaScript: A user complaint or metric says something feels slow
- In real life: Vitals are measured before treatment
- In JavaScript: Use marks, traces, and RUM before changing code
- In real life: A lab test is controlled
- In JavaScript: Lighthouse or a DevTools trace is a repeatable lab run
- In real life: A fitness tracker shows daily life
- In JavaScript: RUM shows real users across devices, networks, and routes
Where the analogy stops: Medical tests have established clinical ranges. Web measurements depend on device, browser, privacy settings, page state, and how you instrument the code.
If you need a refresher on tool basics, revisit Debugging in the browser, Advanced debugging, Observers, Page lifecycle, and the event loop.
Percentiles beat averages for user-facing speed
MEASURE FIRSTPremature optimization often starts with the wrong number. Average response time can look fine while a meaningful slice of users waits. Percentiles answer a different question: “What value did 50%, 75%, or 95% of samples stay under?” Product teams often watch p75 for Core Web Vitals-style decisions and p95 for the slow tail.
const responseTimes = [120, 80, 1000, 160, 240]; function percentile(values, p) { if (values.length === 0) throw new RangeError("percentile needs values"); const sorted = [...values].sort((a, b) => a - b); const rank = Math.ceil((p / 100) * sorted.length); const index = Math.min(sorted.length - 1, Math.max(0, rank - 1)); return sorted[index];} const average = responseTimes.reduce((total, value) => total + value, 0) / responseTimes.length;console.log("average", average);console.log("p50", percentile(responseTimes, 50));console.log("p75", percentile(responseTimes, 75));console.log("p95", percentile(responseTimes, 95));Read it line by line. Line 5 sorts a copy, so the original list stays untouched. Line 6 turns the percentile into a 1-based rank. Line 7 converts that rank to a safe array index. Line 8 returns the value at that index. In this sample, the average is 320 ms, but p75 is 240 ms and p95 is 1000 ms, so the slow tail is impossible to miss.
Step through a real nearest-rank percentile. Watch how one outlier changes p95 more than the average explains.
script
function percentile(values, p) { if (values.length === 0) throw new RangeError("percentile needs values"); const sorted = [...values].sort((a, b) => a - b); const rank = Math.ceil((p / 100) * sorted.length); const index = Math.min(sorted.length - 1, Math.max(0, rank - 1)); return sorted[index];} const average = responseTimes.reduce((total, value) => total + value, 0) / responseTimes.length;console.log("average", average);console.log("p50", percentile(responseTimes, 50));console.log("p75", percentile(responseTimes, 75));console.log("p95", percentile(responseTimes, 95));A page can feel fast if the first useful content appears quickly, even while background work continues. It can also feel slow when the average request is fine but the main thread blocks one important click. Measurement connects the number to the user-visible moment.
The Performance API: clocks, marks, and measures
USER TIMINGUse performance.now() for elapsed time. It is monotonic, high-resolution, and relative to the page time origin. Browsers coarsen precision for security: a normal non-isolated Chrome 154 page in this environment advanced in 0.1 ms steps, while cross-origin isolated pages can expose finer resolution around 5 microseconds. Treat those figures as browser policy, not a promise.
Date.now() is a trap for elapsed timelet wallClock = 1000;const DateNow = () => wallClock; const started = DateNow();wallClock -= 50; // user or OS adjusts the clock backwardconst elapsed = DateNow() - started; console.log(elapsed);console.log("performance.now is monotonic instead");Date.now() reads the wall clock. If the user, operating system, or virtual machine changes that clock, elapsed time can become negative. performance.now() is the right stopwatch. User Timing adds labels: performance.mark() for named points and performance.measure() for spans between them.
function measure(label, fn) { const start = label + ":start"; const end = label + ":end"; performance.mark(start, { detail: { label, phase: "start" } }); const result = fn(); performance.mark(end, { detail: { label, phase: "end" } }); performance.measure(label, { start, end, detail: { label, result } }); performance.clearMarks(start); performance.clearMarks(end); return { result, entries: performance.getEntriesByName(label, "measure") };} const report = measure("hydrate-card", () => "ready");console.log(report.result);console.log(report.entries.at(-1).name);console.log(report.entries.at(-1).duration >= 0);performance.clearMeasures("hydrate-card");Line 4 records a start mark with a detail object. Line 5 runs the work. Line 6 records the end mark. Line 7 creates a measure entry, line 8 and line 9 clear temporary marks, and line 10 returns the original result plus entries from getEntriesByName. Use getEntriesByType("measure") when you want every measure of that type.
Step through a small User Timing helper. It records real marks and measures, then clears the temporary marks.
script
function measure(label, fn) { const start = label + ":start"; const end = label + ":end"; performance.mark(start, { detail: { label, phase: "start" } }); const result = fn(); performance.mark(end, { detail: { label, phase: "end" } }); performance.measure(label, { start, end, detail: { label, result } }); performance.clearMarks(start); performance.clearMarks(end); return { result, entries: performance.getEntriesByName(label, "measure") };} console.log(report.result);console.log(report.entries.at(-1).name);console.log(report.entries.at(-1).duration >= 0);performance.clearMeasures("hydrate-card");function measure(label, fn) { const start = label + ":start"; const end = label + ":end"; performance.mark(start, { detail: { label, phase: "start" } }); const result = fn(); performance.mark(end, { detail: { label, phase: "end" } }); performance.measure(label, { start, end, detail: { label, result } }); performance.clearMarks(start); performance.clearMarks(end); return { result, entries: performance.getEntriesByName(label, "measure") };} const report = measure("hydrate-card", () => "ready");console.log(report.result);console.log(report.entries.at(-1).name);console.log(report.entries.at(-1).duration >= 0);performance.clearMeasures("hydrate-card");Waiting for the browser to run the first measurement.
Run a task to create marks and a measure entry from the live browser Performance API.
measure helper shown in the article. Durations vary by device and browser state.console.time is a quick local probeconsole.time("filter visible products");["keyboard", "mouse", "monitor"].filter(item => item.includes("o"));console.timeEnd("filter visible products");console.time and console.timeEnd are useful scratch tools, but they do not create User Timing entries, do not show up as named spans in a Performance recording, and are easy to leave behind. Promote important probes to marks and measures.
PerformanceObserver: subscribe to timing entries
OBSERVEPerformanceObserver lets code subscribe to entries as the browser creates them. buffered: true asks for earlier entries of that type too, which matters for paint, LCP, navigation, and marks that may have happened before your observer was created.
const types = [ "mark", "measure", "longtask", "long-animation-frame", "navigation", "resource", "paint", "largest-contentful-paint", "event",]; console.log(types.filter(type => PerformanceObserver.supportedEntryTypes.includes(type)).join(", ")); const observer = new PerformanceObserver(list => { for (const entry of list.getEntries()) { console.log(entry.entryType, entry.name); }}); observer.observe({ type: "mark", buffered: true });performance.mark("search-start", { detail: { route: "/products" } });observer.disconnect();Chrome 154 reports support for mark, measure, longtask, long-animation-frame, navigation, resource, paint, largest-contentful-paint, and event among other entries. Long tasks, LoAF, LCP, and event timing are only introduced here; the main-thread and loading-performance lessons use them in depth.
not reportednot reportednot reportednot reportednot reportednot reportednot reportednot reportednot reportedconst nav = performance.getEntriesByType("navigation")[0]; if (nav) { const breakdown = { dns: nav.domainLookupEnd - nav.domainLookupStart, tcp: nav.connectEnd - nav.connectStart, ttfb: nav.responseStart - nav.requestStart, domContentLoaded: nav.domContentLoadedEventEnd - nav.startTime, load: nav.loadEventEnd - nav.startTime, }; console.log(JSON.stringify(breakdown));}| Part | Calculation | What it means |
|---|---|---|
| DNS | domainLookupEnd - domainLookupStart | Time to resolve the host name. |
| TCP | connectEnd - connectStart | Time to create the connection, including TLS in secure contexts. |
| TTFB | responseStart - requestStart | Time from sending the request to the first response byte. |
| DOMContentLoaded | domContentLoadedEventEnd - startTime | When the document is parsed and deferred scripts have run. |
| load | loadEventEnd - startTime | When load-event resources finished and the load event completed. |
The Chrome DevTools Performance panel
TRACEDevTools answers “what happened during this interaction?” Open the Performance panel, use the Live metrics screen for a quick Core Web Vitals-style read, then record runtime performance or record and reload for page load. Use Capture settings to choose CPU and Network throttling. DevTools may recommend presets based on field data.
| Area | How to use it |
|---|---|
| Live metrics screen | Shows Core Web Vitals-style local metrics, optional CrUX field metrics, interactions, layout shifts, and environment settings. |
| Main track | Shows tasks on the main thread. Long tasks are marked with red triangles when they exceed the long-task threshold. |
| Bottom-up | Starts from the heaviest leaf functions so self time stands out. |
| Call tree | Starts from roots and expands into callees so total time by path stands out. |
| Event log | Lists trace events in chronological order for exact sequencing. |
| Insights sidebar | Surfaces guided findings also used by Lighthouse-style performance insights. |
In the recording, the Main track shows tasks on the main thread. Long tasks get red triangle markers. User Timing marks and measures appear on the Timings track so your labels line up with script, style, layout, paint, screenshots, requests, and frames. The lower tabs help you change the question: Bottom-up asks “which leaves cost the most?”, Call tree asks “which path cost the most?”, and Event log asks “what happened in exact order?” The Insights tab or sidebar highlights issues such as LCP and INP subparts.
- Reproduce one slow interaction.
- Record with realistic CPU and network settings.
- Look for one long task or wide flame-chart path.
- Confirm the same issue with a second run or field data before refactoring.
Flame charts: x is time, y is stack depth
CALL STACKSA flame chart is a stacked timeline. The horizontal position is when a function was on the stack. Width is total duration. Vertical nesting is call stack depth. A parent includes its children, so total time is not the same as self time. Self time is the part of a frame not covered by child frames.
function profileCheckout() { return profile("checkout", () => { parseCart(); priceItems(); renderSummary(); });} function priceItems() { readPrices(); applyDiscounts();}Selected applyDiscounts: total 9 ms, self 9 ms. The widest frame is checkout; the largest self time is applyDiscounts.
Step through the tiny chart before opening a huge trace. checkout is widest because it contains everything. Inside it, priceItems contains readPrices and applyDiscounts. applyDiscounts has the largest self time because it does the most direct work in this recorded tree. A real browser trace is sampled, so tiny functions may appear or disappear depending on sampling and inlining.
Real user monitoring: field data complements lab data
RUMLab data is controlled: Lighthouse, WebPageTest, Playwright traces, and local DevTools recordings. Field data comes from real users. Chrome User Experience Report (CrUX) provides public field data for eligible origins and URLs. Real user monitoring is your own sampled telemetry, sliced by route, device class, release, connection, and privacy-safe dimensions.
- A Lighthouse run on one machine with CPU and network throttling.
- CrUX p75 LCP for your origin on mobile users over the last 28 days.
- Your app sends INP and route metrics with
navigator.sendBeaconon page hide. - A local tinybench or mitata suite compares two string-formatting functions.
- A DevTools recording of a checkout click on a test device.
- A loop runs two implementations 10,000 times and reports median sample time.
Sort each measurement by what kind of evidence it gives you. Use both kinds before a high-risk optimization.
web-vitals RUM sketchJavaScriptimport { onCLS, onINP, onLCP, onTTFB } from "web-vitals"; function sendMetric(metric) { navigator.sendBeacon("/rum", JSON.stringify({ name: metric.name, value: metric.value, id: metric.id, }));} onCLS(sendMetric);onINP(sendMetric);onLCP(sendMetric);onTTFB(sendMetric);The web-vitals package is not installed in this app, so the import above is a production sketch, not a runnable lesson example. The current API exposes callbacks such as onCLS, onINP, onLCP, and onTTFB. The callback receives a metric object with fields such as name, value, id, delta, and entries.
const queue = []; function rememberMetric(metric) { queue.push({ name: metric.name, value: metric.value, path: location.pathname });} document.addEventListener("visibilitychange", () => { if (document.visibilityState !== "hidden" || queue.length === 0) return; const payload = JSON.stringify(queue.splice(0)); navigator.sendBeacon("/rum", payload);});Use navigator.sendBeacon during visibilitychange because users close tabs and navigate away. Sample traffic, remove personal data, avoid raw URLs with secrets, and document retention. RUM is evidence, not surveillance.
Benchmarking pitfalls: make the test harder to fool
BENCHMARKSMicrobenchmarks are useful when the question is tiny and isolated, but they are easy to fool. Watch for JIT warm-up, unused results being optimized away, timer resolution, variance, outliers, DevTools overhead, garbage collection, and code that does not resemble production. A macro benchmark or real trace often matters more than a faster loop.
function median(values) { const sorted = [...values].sort((a, b) => a - b); const middle = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[middle] : (sorted[middle - 1] + sorted[middle]) / 2;} function benchmark(label, fn, { warmup = 20, samples = 7, iterations = 200 } = {}) { let sink = 0; for (let i = 0; i < warmup; i += 1) sink += fn(); const durations = []; for (let sample = 0; sample < samples; sample += 1) { const start = performance.now(); for (let i = 0; i < iterations; i += 1) sink += fn(); durations.push(performance.now() - start); } return { label, median: median(durations), samples: durations.length, sink };} const report = benchmark("math", () => Math.sqrt(144));console.log(report.label);console.log(report.samples);console.log(report.median >= 0);console.log(report.sink > 0);Step through a safer benchmark runner: warm up, repeat, consume results, and report a median instead of one lucky number.
script
function median(values) { const sorted = [...values].sort((a, b) => a - b); const middle = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[middle] : (sorted[middle - 1] + sorted[middle]) / 2;} function benchmark(label, fn, { warmup = 20, samples = 7, iterations = 200 } = {}) { let sink = 0; for (let i = 0; i < warmup; i += 1) sink += fn(); const durations = []; for (let sample = 0; sample < samples; sample += 1) { const start = performance.now(); for (let i = 0; i < iterations; i += 1) sink += fn(); durations.push(performance.now() - start); } return { label, median: median(durations), samples: durations.length, sink };} const report = benchmark("math", () => Math.sqrt(144));console.log(report.label);console.log(report.samples);console.log(report.median >= 0);function median(values) { const sorted = [...values].sort((a, b) => a - b); const middle = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[middle] : (sorted[middle - 1] + sorted[middle]) / 2;} function benchmark(label, fn, { warmup = 20, samples = 7, iterations = 200 } = {}) { let sink = 0; for (let i = 0; i < warmup; i += 1) sink += fn(); const durations = []; for (let sample = 0; sample < samples; sample += 1) { const start = performance.now(); for (let i = 0; i < iterations; i += 1) sink += fn(); durations.push(performance.now() - start); } return { label, median: median(durations), samples: durations.length, sink };} const report = benchmark("math", () => Math.sqrt(144));console.log(report.label);console.log(report.samples);console.log(report.median >= 0);console.log(report.sink > 0);Waiting for the browser to run the benchmark.
The exact milliseconds are intentionally not the lesson. The shape is: one cold duration versus multiple samples and a median after warm-up.
Libraries such as tinybench and mitata help with warm-up, repeated samples, and reports. They do not remove your responsibility to choose production-like inputs, run with DevTools closed when appropriate, compare medians and variance, and confirm the result in the real interaction.
Common misconceptions
“The average is the user experience.”
Averages flatten the tail. Use p50, p75, and p95 to see typical, target, and slow-user experiences.
“A faster microbenchmark always makes the app faster.”
Not if the code is not on the hot path, the benchmark was optimized away, or rendering and network dominate.
“DevTools open is the same as production.”
DevTools can add overhead and change sampling. Use it to diagnose, then validate in production-like conditions.
“RUM replaces lab traces.”
RUM tells you where users hurt. Lab traces help explain why and let you verify a fix before release.
“A mark is automatically cleaned up.”
Marks stay in the performance timeline until cleared or the document goes away. Clear temporary marks after measuring.
| Tool | Best question | Watch out |
|---|---|---|
Date.now() | Wall-clock time since the Unix epoch. | Logging timestamps, not measuring elapsed work; the clock can jump backward or forward. |
performance.now() | Monotonic elapsed time from the page's time origin. | Short measurements; high resolution is coarsened for privacy, commonly around 100 microseconds unless cross-origin isolated. |
| User Timing | Named mark and measure entries with optional detail. | Timing your own app phases and showing them in DevTools Timings. |
| DevTools profile | A recorded trace of tasks, network, frames, screenshots, and sampled JavaScript stacks. | Finding where time went in a real interaction. |
| RUM | Metrics collected from real users' browsers. | Watching p75/p95 performance by route, device class, region, and release. |
Practice exercises
5 EXERCISESUse nearest-rank percentiles. What is p75 for the sample?
const samples = [120, 80, 1000, 160, 240];
// sorted: 80, 120, 160, 240, 1000
// p75 rank: Math.ceil(0.75 * 5)console.log(240);p75 is 240 because four out of five sorted samples are at or below 240 ms.
What elapsed value does the snippet print first?
let wallClock = 1000;
const DateNow = () => wallClock;
const started = DateNow();
wallClock -= 50; // user or OS adjusts the clock backward
const elapsed = DateNow() - started;
console.log(elapsed);
console.log("performance.now is monotonic instead");The first log is -50. Use performance.now() for elapsed time because it is monotonic.
Which Performance API method removes the temporary marks after a measure is created?
performance.clearMarks(start);
performance.clearMarks(end);Clear marks after creating a measure so later runs do not accidentally reuse stale names.
Which function has the highest self time?
checkout: 25 ms total, 0 ms self
parseCart: 4 ms total, 4 ms self
priceItems: 14 ms total, 0 ms self
readPrices: 5 ms total, 5 ms self
applyDiscounts: 9 ms total, 9 ms self
renderSummary: 7 ms total, 7 ms selfapplyDiscounts has 9 ms self time, the largest direct work in the recorded tree.
Which statistic should the better benchmark report instead of one cold duration?
// Fill the summary word: warm up, keep a sink, run many samples, report the ____.return { median: median(durations), samples: durations.length, sink };Report the median and sample count after warm-up, and keep a sink so results are used.
Quiz: check your understanding
8 QUESTIONSAnswer from the evidence: what was measured, where it came from, and what could fool it?
Question 1 of 8Why does this lesson prefer p75 and p95 over an average for user-facing speed?
Choose an answer to see the explanation.
Question 2 of 8What does the percentile snippet print for p75?
Read the code, then predictconst samples = [120, 80, 1000, 160, 240]; const sorted = [...samples].sort((a, b) => a - b); const rank = Math.ceil((75 / 100) * sorted.length); console.log(sorted[rank - 1]);Choose an answer to see the explanation.
Question 3 of 8What makes
performance.now()safer for elapsed-time measurement thanDate.now()?Choose an answer to see the explanation.
Question 4 of 8What does this
Date.now()pitfall print first?Read the code, then predictlet wallClock = 1000; const DateNow = () => wallClock; const started = DateNow(); wallClock -= 50; console.log(DateNow() - started);Choose an answer to see the explanation.
Question 5 of 8Where do User Timing marks appear in a Chrome DevTools Performance recording?
Choose an answer to see the explanation.
Question 6 of 8How do you read a flame chart rectangle?
Choose an answer to see the explanation.
Question 7 of 8Which statement separates RUM from a Lighthouse lab run?
Choose an answer to see the explanation.
Question 8 of 8Which benchmark result should you trust most?
Choose an answer to see the explanation.
Key takeaways
- Measure the user-visible problem first; premature optimization can improve the wrong thing.
- Use percentiles, especially p75 and p95, instead of trusting averages.
performance.now(), marks, measures, observers, and navigation timing give browser-grounded evidence.- Chrome DevTools Performance recordings connect tasks, marks, screenshots, flame charts, and Insights.
- RUM prioritizes real-user pain; lab traces explain causes; benchmarks need warm-up, used results, many samples, and medians.
Remember the one-liner.
Measure first, choose the bottleneck second, optimize last, and verify with the same measurement.
Up next: Keeping the main thread responsive.