CPU profiling in depth
Learn how CPU profilers sample stacks, read flame charts and bottom-up views, and use Node and Linux profiling tools carefully.
- 01Read sampled evidenceExplain what a stack sample estimates and why its counts are not exact stopwatch timings.
- 02Find expensive workUse flame-chart width and bottom-up self time to choose the next function to inspect.
- 03Choose a toolCollect a Node profile, recognize optimized frames, and know when runtime call stats help.
What CPU profiles answer
A CPU profile is evidence about where a program spent processor time during one run. It helps when a page feels slow, a request uses too much CPU, or an animation drops frames. It does not tell you what to change until you read the call path and reproduce the work.
CPU profiling collects observations about active code while a program runs. A profile groups those observations so you can find hot paths before optimizing.
The first rule is simple: look for a repeated cost, not a scary function name. Our tiny order has one long loop and one short return. The loop should be hot because it remains active for more of the run.
function makeTea() { let heat = 0; for (let cup = 0; cup < 30; cup += 1) heat += cup; return heat;} function serve() { return "tea";} function handleOrder() { return `${serve()}: ${makeTea()}`;} console.log(handleOrder());Line 1 starts makeTea. Line 3 repeats thirty times and returns 435. Line 7 starts serve and returns quickly. Line 11 calls both functions, so line 14 prints tea: 435.
In a real application the loop might be parsing a response, drawing a chart, or filtering a large list. The names change. The profile question stays the same: which active path appears often enough to explain the slow run?
A profile is a map of one trip through the program. It cannot decide whether a result is useful, and it cannot replace a user report. It gives you a place to look before changing code that only looks suspicious.
Start with the action that feels slow. For example, click the same button with the same cart size three times. A repeatable action makes a hot path easier to separate from background work.
Sampling and instrumenting
A sampling profiler takes stack snapshots at intervals. It is like taking a photo of a busy home kitchen every second. If the cook is at the stove in 80 of 100 photos, cooking took roughly most of the observed time.
Take photos of a busy kitchen at regular moments. The cook who appears in most photos is probably doing the work that occupies the kitchen.
- In real life: A photo catches the cook at the stove
- In JavaScript: A snapshot catches a function on the stack
- In real life: Many stove photos
- In JavaScript: More samples include that function
- In real life: A quick action between photos
- In JavaScript: Short work can be missed
- In real life: A stopwatch on every step
- In JavaScript: Instrumentation records exact events
Where the analogy stops: Photos estimate where time lands; they do not measure each individual call exactly.
An instrumenting profiler puts a measurement around every enter and exit. It can count exact calls, but those measurements also slow the program. Sampling usually gives a lower-overhead picture of a production-shaped run.
| Tool | How it observes | What it is good for |
|---|---|---|
| Sampling profiler | Takes occasional stack photos | Low overhead estimate of where time lands |
| Instrumenting profiler | Records each entered and exited function | Exact events, with more observer overhead |
| Flame chart | Merged call stacks, width by sample count | Which paths occupied observed time |
| Bottom-up view | Leaf self time sorted largest first | Which function was busy itself |
| Self time | A sampled function is the active leaf | Work performed by its children |
| Total time | A sampled function appears anywhere on the stack | Proof that its own body is expensive |
| Runtime call stats | Engine categories in a trace | A replacement for application call paths |
let calls = 0;
function measure(name, work) {
calls += 1;
const result = work();
console.log(name, result);
return result;
}
measure("serve", () => "tea");
console.log(calls);Line 1 creates a counter. Line 4 runs a piece of work while line 5 prints its name and result. Line 9 measures one call to serve, then line 10 prints 1. This is intentionally tiny: a real instrumenting profiler records many more events.
The important cost is the measurement itself. The wrapper has to call a function, update a counter, and log a value for every observed operation. That makes it useful for exact call information, but less faithful when the program is already near its CPU limit.
Neither method is morally better. Use an instrumented trace when exact events matter. Use a sampler when you need the run to stay close to normal and want to find the big shapes first.
A small sampling model
This model receives scripted stacks rather than pausing a real engine. That makes every replay stable enough to inspect. It counts the final frame as self time and counts every frame in a stack as total time.
Replay of instrumented lesson code. It demonstrates sampling; it is not an engine debugger.
script
["handleOrder", "makeTea"], ["handleOrder", "makeTea"], ["handleOrder", "serve"],]; const self = {};for (const stack of stacks) { const leaf = stack.at(-1); self[leaf] = (self[leaf] ?? 0) + 1;}console.log(self.makeTea);Line 1 creates three snapshots. Line 7 calls the sampler. Two snapshots end at makeTea, so its self count is two. The model does not claim two samples equal two milliseconds; they are observations.
Change the interval later in the playground. Fewer photos can still show the same broad story, but a small run can also change noticeably. That uncertainty is why profiles guide investigation instead of proving a fixed percentage.
const stacks = [
["handleOrder", "makeTea"],
["handleOrder", "serve"],
];
const leafs = stacks.map((stack) => stack.at(-1));
console.log(leafs.filter((name) => name === "makeTea").length);Line 1 creates two observations. Line 5 keeps the final frame from each stack. Line 6 counts one makeTea leaf and prints 1. With more samples, that count becomes an estimate of the share of observed time.
A sampler may catch different moments when its interval changes. That is normal. Use a longer recording when a small difference would change your decision, and trust repeated large patterns more than one narrow bar.
Flame charts merge call paths
A flame chart merges identical stacks into a tree. A bar represents one function in one call path. Its width is the number of samples that included that frame. It is not a literal timeline of every call.
Replay of a real tree-building function. The bars in the playground are a visualization of its sample counts.
script
["handleOrder", "makeTea"], ["handleOrder", "makeTea"], ["handleOrder", "serve"],]; function buildFlameTree(stacks) { return { count: stacks.length };}const tree = buildFlameTree(stacks);console.log(tree.count);Line 1 names the same three stacks. Line 7 merges their prefixes. The shared handleOrder bar has width three because all three observations passed through it. The children divide that width between makeTea and serve.
const stacks = [
["handleOrder", "makeTea"],
["handleOrder", "makeTea"],
["handleOrder", "serve"],
];
const self = {};
for (const stack of stacks) {
const leaf = stack.at(-1);
self[leaf] = (self[leaf] ?? 0) + 1;
}
console.log(self.makeTea);This interval keeps 6 observations. makeTea has 5 leaf samples, estimated at 83% of observed leaves.
The control changes one input: the snapshot interval. Reset restores one-tick sampling. When an estimate changes, ask whether the broad hot path remains, then take a longer profile before making a decision.
Self time and total time use the same snapshots. A makeTea leaf gets self time when it is the last frame in a photo. handleOrder gets total time because every photo passed through it on the way to a leaf.
That distinction prevents a common mistake. The caller can have the widest total bar while doing almost no work itself. Open the caller, then follow the wide child until a leaf keeps appearing.
Bottom-up views start at the leaf
A bottom-up view sorts functions by self time. It begins with the active leaf, then shows callers that led there. This is useful when one helper is expensive but its callers are many and would dominate a top-down view.
const rows = [["makeTea", 5], ["serve", 1]];
console.log(rows.sort((a, b) => b[1] - a[1])[0][0]);Line 1 creates two self-time rows. Line 2 sorts the larger count first and prints makeTea. That output matches the sampling model: five self samples beat one self sample.
Read both views together. A flame chart shows context: who called the work. A bottom-up table says which leaf was busy itself. A wide caller can be innocent if all of its time belongs to one child.
Suppose three screens call makeTea. A flame chart keeps those three paths separate so you can see context. Bottom-up combines the leaf samples, making the shared helper obvious even when no individual caller looks dramatic.
Then return to the callers. A hot helper may need a cache, fewer inputs, or less frequent calls. The bottom-up list chooses the function to understand; it does not prescribe an optimization.
Node and Linux tools
Node has two useful CPU profile routes. --prof writes V8 tick data that --prof-process summarizes. --cpu-prof writes a DevTools-readable .cpuprofile file, which Chrome DevTools can open.
# Node's tick profiler
node --prof app.js
node --prof-process isolate-*.log
# Chrome DevTools profile file
node --cpu-prof --cpu-prof-dir=profiles app.js
# Linux only: native stacks for perf
perf record -g node --perf-basic-prof app.jsThe first two lines run and process V8 ticks. The next command writes a file into a chosen directory. The final command is Linux-only: perf record -g records call stacks while Node exposes symbols with --perf-basic-prof.
The test for this lesson runs a tiny child process with --cpu-prof, reads its JSON file, and checks only stable facts: a unique hot function is present in nodes and samples is not empty. It never asserts timing or a sample count.
For node --prof, run the workload once, then pass the generated isolate log to node --prof-process. Its text summary is useful when you want V8's view of ticks. Keep the raw log beside the command that produced it so the report remains reproducible.
For node --cpu-prof, open the generated file in Chrome DevTools Performance panel. The file contains nodes and samples, so it can show source names, callers, and a timeline-like view. It is usually the friendliest first format for a JavaScript team.
Inlined and optimized frames
Optimizing compilers can inline a small function into its caller. That means the machine may execute serve as part of handleOrder, without an ordinary call boundary. JavaScript output must stay the same.
function serve() { return "tea"; }
function handleOrder() { return serve(); }
// After optimization, an engine may inline serve into handleOrder.
// DevTools can reconstruct an inlined serve frame from source positions.
console.log(handleOrder());Line 1 returns tea. Line 2 calls it. Line 6 prints tea. After optimization, DevTools may reconstruct a serve frame from source positions, but its timing can shift because the physical work was merged.
Do not treat a reconstructed inlined frame as a lie. Treat it as a source-level explanation. Compare profiles before and after a code change, look at the surrounding path, and avoid assuming every visible frame is one machine-level call.
Here is a useful mental model. Before optimization, the machine can enter serve and return to handleOrder. After inlining, the instructions for returning tea can sit inside the caller, so the measured time may move up to the caller.
DevTools still has source positions and optimization metadata when available. It can display an inlined serve frame to help you reason in JavaScript names. Treat that display as a helpful reconstruction, not a promise about one physical stack frame.
Runtime call stats
Runtime call stats are V8 counters that group time spent in engine work. In Chrome tracing tools they can expose categories such as parsing, compiling, garbage collection, and inline-cache misses.
// Chrome tracing can expose V8 runtime call stats.
// They group engine work such as parsing, compiling, GC, and IC misses.
// Use them to ask where engine time went, not to time one tiny call.
console.log("record a trace, then inspect RuntimeCallStats");Line 1 is a comment, not a browser API call. Lines 2 and 3 name the kind of engine work these counters group. Line 4 prints a reminder: record a trace and inspect the RuntimeCallStats view.
Use this when application frames alone do not explain the cost. It is a diagnostic category view, not a promise that every engine version exposes identical counters or names.
For example, a profile can show a short application function that triggers a lot of parsing or garbage collection. Runtime call stats move the question one level lower: is the engine spending time compiling new code, collecting temporary objects, or handling inline-cache misses?
Keep this tool brief and targeted. First find the slow user action and its JavaScript path. Then inspect tracing categories only when the application frames leave a meaningful part of the cost unexplained.
Profiling a real problem
Start with a real action: typing in search, opening a cart, or loading a report. Reproduce it with realistic data. Record one baseline profile, locate a repeatable hot path, then change one idea and record again.
- Profile a representative run, not an empty page.
- Use a flame chart to find context and bottom-up self time to find busy leaves.
- Keep the result of work alive so the optimizer cannot remove a benchmark.
- Compare before and after profiles under similar conditions.
function filterOrders(orders, query) {
return orders.filter((order) => order.name.includes(query));
}
const orders = [{ name: "tea" }, { name: "toast" }, { name: "tea cake" }];
const visible = filterOrders(orders, "tea");
console.log(visible.length);Line 1 defines a filter used by a search page. Line 5 creates three visible orders. Line 6 filters for tea, and line 7 prints 2. The code is small on purpose: the same shape becomes expensive when a real page repeats it over a large list during every keystroke.
Find the hot function in a slow page
First, open the slow page with the normal amount of data. Start a recording, type one search term, and stop the recording after the results settle. Do not record an empty state if customers report slowness with a full list.
Second, use the flame chart to follow the wide path from the input event to a repeated function. Third, open bottom-up and look for the leaf with the largest self count. If that leaf is filterOrders, inspect how often it runs and how many orders it scans.
Finally, make one change, such as filtering only after a short pause or avoiding a duplicate filter. Record the identical action again. Keep the change only when the repeated hot path shrinks without breaking the result.
- Take one stack photo every few ticks
- Record every function enter and exit
- Merge repeated caller paths into bars
- Sort leaf counts from largest to smallest
- Choose a periodic observation interval
- Add measurement around every call
Place each card by how it collects evidence or presents it.
For user-facing work, pair CPU evidence with measuring performance and main-thread. For compiler behavior, return to optimizing compilers.
Common mistakes
A profile is strongest when it changes a question into a smaller question. “Why is the page slow?” becomes “Why does this leaf appear so often for this action?” That keeps guesses out of the investigation.
- “The widest frame is always the bug.” It may be a caller whose child owns the self time.
- “One profile is a benchmark.” It is one observation of one workload.
- “Inlining removes source context.” Tools can reconstruct source frames, though costs can move.
- “A sample count is exact time.” It is an estimate shaped by the interval and run.
| Term | What it says | What it does not say |
|---|---|---|
| A wide bar | Many samples include that frame | A promise that the source line is slow |
| Self time | Samples where a function is the leaf | Time including every child call |
| An inlined frame | A source-level frame reconstructed by tools | Proof the machine executed a normal call |
| A short missing function | Work that may fall between sample ticks | Proof that it never ran |
When two runs disagree, check the workload, the machine state, and the recording length before changing the code. A new profile is evidence to compare, not a score to win.
Practice exercises
Run the small model and type the printed count.
const stacks = [["handleOrder", "makeTea"], ["handleOrder", "makeTea"], ["handleOrder", "serve"]];
console.log(stacks.filter((stack) => stack.at(-1) === "makeTea").length);It prints 2, because two of the three leaf frames are makeTea.
The report has five makeTea leaf samples and one serve leaf sample. Which function should you inspect first?
makeTea is hottest by self samples. It is the leaf in five observations.
Choose the flag that writes a DevTools-readable CPU profile file.
Use --cpu-prof. Node writes a .cpuprofile file that Chrome DevTools can open.
What does the inlining model print?
function serve() { return "tea"; }
function handleOrder() { return serve(); }
console.log(handleOrder());It prints tea. Optimization can change machine frames, not this JavaScript result.
A checkout page feels slow with a large cart. What should you do before changing its rendering code?
Record a profile first. Then inspect the repeatable hot path and compare a focused change against the same action.
Name one kind of engine work that runtime call stats can group.
Runtime call stats can group parsing, compiling, garbage collection, and inline-cache miss work.
Check your understanding
Use the same rule for every question: first identify what the tool observed, then identify what the view groups.
Question 1 of 8What is a sampling profiler doing?
Choose an answer to see the explanation.
Question 2 of 8What does this print?
Read the code, then predictconst stacks = [["handleOrder", "makeTea"], ["handleOrder", "makeTea"], ["handleOrder", "serve"]]; console.log(stacks.filter((stack) => stack.at(-1) === "makeTea").length);Choose an answer to see the explanation.
Question 3 of 8What does a flame-chart bar width represent?
Choose an answer to see the explanation.
Question 4 of 8Why inspect bottom-up self time?
Choose an answer to see the explanation.
Question 5 of 8Which Node flag writes a
.cpuprofilefile?Choose an answer to see the explanation.
Question 6 of 8What may happen after
serveis inlined intohandleOrder?Choose an answer to see the explanation.
Question 7 of 8What are runtime call stats useful for?
Choose an answer to see the explanation.
Question 8 of 8A function appears in every stack but is never the leaf. What does that show?
Choose an answer to see the explanation.
Key takeaways
- Sampling takes periodic stack photos and estimates where execution was observed.
- Instrumentation records exact events but adds overhead.
- Flame charts show merged call context; bottom-up views sort busy leaves.
- Use `--cpu-prof` for a DevTools-readable Node profile and Linux perf when native stacks help.
- Inlining can change physical frames while tools preserve source context.
- Runtime call stats separate engine work from application work.
Remember the one-liner.
Profile the real action, then follow the repeated path before optimizing.
Coming next: Benchmarking without fooling yourself.