cf.completefrontendCode editorOpen lab
THE JAVASCRIPT FIELD GUIDE

CPU profiling in depth

Learn how CPU profilers sample stacks, read flame charts and bottom-up views, and use Node and Linux profiling tools carefully.

By the end, you can
  • 01
    Read sampled evidenceExplain what a stack sample estimates and why its counts are not exact stopwatch timings.
  • 02
    Find expensive workUse flame-chart width and bottom-up self time to choose the next function to inspect.
  • 03
    Choose a toolCollect a Node profile, recognize optimized frames, and know when runtime call stats help.

What CPU profiles answer

A CPU profile is evidence about where a program spent processor time during one run. It helps when a page feels slow, a request uses too much CPU, or an animation drops frames. It does not tell you what to change until you read the call path and reproduce the work.

Definition

CPU profiling collects observations about active code while a program runs. A profile groups those observations so you can find hot paths before optimizing.

The first rule is simple: look for a repeated cost, not a scary function name. Our tiny order has one long loop and one short return. The loop should be hot because it remains active for more of the run.

A tiny orderPop out in the code editor (opens in a new tab)JavaScript
function makeTea() {  let heat = 0;  for (let cup = 0; cup < 30; cup += 1) heat += cup;  return heat;} function serve() {  return "tea";} function handleOrder() {  return `${serve()}: ${makeTea()}`;} console.log(handleOrder());

Line 1 starts makeTea. Line 3 repeats thirty times and returns 435. Line 7 starts serve and returns quickly. Line 11 calls both functions, so line 14 prints tea: 435.

In a real application the loop might be parsing a response, drawing a chart, or filtering a large list. The names change. The profile question stays the same: which active path appears often enough to explain the slow run?

A profile is a map of one trip through the program. It cannot decide whether a result is useful, and it cannot replace a user report. It gives you a place to look before changing code that only looks suspicious.

Start with the action that feels slow. For example, click the same button with the same cart size three times. A repeatable action makes a hot path easier to separate from background work.

Sampling and instrumenting

A sampling profiler takes stack snapshots at intervals. It is like taking a photo of a busy home kitchen every second. If the cook is at the stove in 80 of 100 photos, cooking took roughly most of the observed time.

Real-life analogyKitchen photos

Take photos of a busy kitchen at regular moments. The cook who appears in most photos is probably doing the work that occupies the kitchen.

In real life: A photo catches the cook at the stove
In JavaScript: A snapshot catches a function on the stack
In real life: Many stove photos
In JavaScript: More samples include that function
In real life: A quick action between photos
In JavaScript: Short work can be missed
In real life: A stopwatch on every step
In JavaScript: Instrumentation records exact events

Where the analogy stops: Photos estimate where time lands; they do not measure each individual call exactly.

An instrumenting profiler puts a measurement around every enter and exit. It can count exact calls, but those measurements also slow the program. Sampling usually gives a lower-overhead picture of a production-shaped run.

Ways to collect CPU evidence
ToolHow it observesWhat it is good for
Sampling profilerTakes occasional stack photosLow overhead estimate of where time lands
Instrumenting profilerRecords each entered and exited functionExact events, with more observer overhead
Flame chartMerged call stacks, width by sample countWhich paths occupied observed time
Bottom-up viewLeaf self time sorted largest firstWhich function was busy itself
Self timeA sampled function is the active leafWork performed by its children
Total timeA sampled function appears anywhere on the stackProof that its own body is expensive
Runtime call statsEngine categories in a traceA replacement for application call paths
A tiny instrumenting modelPop out in the code editor (opens in a new tab)JavaScript
let calls = 0;

function measure(name, work) {
  calls += 1;
  const result = work();
  console.log(name, result);
  return result;
}

measure("serve", () => "tea");
console.log(calls);

Line 1 creates a counter. Line 4 runs a piece of work while line 5 prints its name and result. Line 9 measures one call to serve, then line 10 prints 1. This is intentionally tiny: a real instrumenting profiler records many more events.

The important cost is the measurement itself. The wrapper has to call a function, update a counter, and log a value for every observed operation. That makes it useful for exact call information, but less faithful when the program is already near its CPU limit.

Neither method is morally better. Use an instrumented trace when exact events matter. Use a sampler when you need the run to stay close to normal and want to find the big shapes first.

A small sampling model

This model receives scripted stacks rather than pausing a real engine. That makes every replay stable enough to inspect. It counts the final frame as self time and counts every frame in a stack as total time.

Step through the sampling model
Step 0 of 4Ready
Your turn: follow the blue line

Replay of instrumented lesson code. It demonstrates sampling; it is not an engine debugger.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
  ["handleOrder", "makeTea"],  ["handleOrder", "makeTea"],  ["handleOrder", "serve"],]; const self = {};for (const stack of stacks) {  const leaf = stack.at(-1);  self[leaf] = (self[leaf] ?? 0) + 1;}console.log(self.makeTea);
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

Line 1 creates three snapshots. Line 7 calls the sampler. Two snapshots end at makeTea, so its self count is two. The model does not claim two samples equal two milliseconds; they are observations.

Change the interval later in the playground. Fewer photos can still show the same broad story, but a small run can also change noticeably. That uncertainty is why profiles guide investigation instead of proving a fixed percentage.

Count a sampled percentagePop out in the code editor (opens in a new tab)JavaScript
const stacks = [
  ["handleOrder", "makeTea"],
  ["handleOrder", "serve"],
];
const leafs = stacks.map((stack) => stack.at(-1));
console.log(leafs.filter((name) => name === "makeTea").length);

Line 1 creates two observations. Line 5 keeps the final frame from each stack. Line 6 counts one makeTea leaf and prints 1. With more samples, that count becomes an estimate of the share of observed time.

A sampler may catch different moments when its interval changes. That is normal. Use a longer recording when a small difference would change your decision, and trust repeated large patterns more than one narrow bar.

Flame charts merge call paths

A flame chart merges identical stacks into a tree. A bar represents one function in one call path. Its width is the number of samples that included that frame. It is not a literal timeline of every call.

Step through merged flame stacks
Step 0 of 4Ready
Your turn: follow the blue line

Replay of a real tree-building function. The bars in the playground are a visualization of its sample counts.

Running in
  1. script
Next: line 1
Click the blue line to take the next stepPop out in the code editor (opens in a new tab)JavaScript
  ["handleOrder", "makeTea"],  ["handleOrder", "makeTea"],  ["handleOrder", "serve"],]; function buildFlameTree(stacks) {  return { count: stacks.length };}const tree = buildFlameTree(stacks);console.log(tree.count);
CallStoreChangeResultRun = next line. Ran = already executed.
Recent returnsNothing yet. Start with the blue line.
A guided replay recorded from real JavaScript calls, not an engine debugger. Step follows executed statements; Back reviews a snapshot. Reset starts a fresh run.

Line 1 names the same three stacks. Line 7 merges their prefixes. The shared handleOrder bar has width three because all three observations passed through it. The children divide that width between makeTea and serve.

Playground: change the sampling interval
Sampling model inputJavaScript
const stacks = [
  ["handleOrder", "makeTea"],
  ["handleOrder", "makeTea"],
  ["handleOrder", "serve"],
];

const self = {};
for (const stack of stacks) {
  const leaf = stack.at(-1);
  self[leaf] = (self[leaf] ?? 0) + 1;
}
console.log(self.makeTea);
Merged stacks
handleOrder
makeTea: 5 samples
serve: 1 samples
makeTea: 5 self samples
serve: 1 self samples
Try it yourself

This interval keeps 6 observations. makeTea has 5 leaf samples, estimated at 83% of observed leaves.

A deterministic teaching model. Real samplers observe a running program and can miss short work between snapshots.

The control changes one input: the snapshot interval. Reset restores one-tick sampling. When an estimate changes, ask whether the broad hot path remains, then take a longer profile before making a decision.

Self time and total time use the same snapshots. A makeTea leaf gets self time when it is the last frame in a photo. handleOrder gets total time because every photo passed through it on the way to a leaf.

That distinction prevents a common mistake. The caller can have the widest total bar while doing almost no work itself. Open the caller, then follow the wide child until a leaf keeps appearing.

Bottom-up views start at the leaf

A bottom-up view sorts functions by self time. It begins with the active leaf, then shows callers that led there. This is useful when one helper is expensive but its callers are many and would dominate a top-down view.

Sort the leaf countsPop out in the code editor (opens in a new tab)JavaScript
const rows = [["makeTea", 5], ["serve", 1]];
console.log(rows.sort((a, b) => b[1] - a[1])[0][0]);

Line 1 creates two self-time rows. Line 2 sorts the larger count first and prints makeTea. That output matches the sampling model: five self samples beat one self sample.

Read both views together. A flame chart shows context: who called the work. A bottom-up table says which leaf was busy itself. A wide caller can be innocent if all of its time belongs to one child.

Suppose three screens call makeTea. A flame chart keeps those three paths separate so you can see context. Bottom-up combines the leaf samples, making the shared helper obvious even when no individual caller looks dramatic.

Then return to the callers. A hot helper may need a cache, fewer inputs, or less frequent calls. The bottom-up list chooses the function to understand; it does not prescribe an optimization.

Node and Linux tools

Node has two useful CPU profile routes. --prof writes V8 tick data that --prof-process summarizes. --cpu-prof writes a DevTools-readable .cpuprofile file, which Chrome DevTools can open.

Node and Linux commandsShell
# Node's tick profiler
node --prof app.js
node --prof-process isolate-*.log

# Chrome DevTools profile file
node --cpu-prof --cpu-prof-dir=profiles app.js

# Linux only: native stacks for perf
perf record -g node --perf-basic-prof app.js

The first two lines run and process V8 ticks. The next command writes a file into a chosen directory. The final command is Linux-only: perf record -g records call stacks while Node exposes symbols with --perf-basic-prof.

The test for this lesson runs a tiny child process with --cpu-prof, reads its JSON file, and checks only stable facts: a unique hot function is present in nodes and samples is not empty. It never asserts timing or a sample count.

For node --prof, run the workload once, then pass the generated isolate log to node --prof-process. Its text summary is useful when you want V8's view of ticks. Keep the raw log beside the command that produced it so the report remains reproducible.

For node --cpu-prof, open the generated file in Chrome DevTools Performance panel. The file contains nodes and samples, so it can show source names, callers, and a timeline-like view. It is usually the friendliest first format for a JavaScript team.

Inlined and optimized frames

Optimizing compilers can inline a small function into its caller. That means the machine may execute serve as part of handleOrder, without an ordinary call boundary. JavaScript output must stay the same.

An inlining modelPop out in the code editor (opens in a new tab)JavaScript
function serve() { return "tea"; }
function handleOrder() { return serve(); }

// After optimization, an engine may inline serve into handleOrder.
// DevTools can reconstruct an inlined serve frame from source positions.
console.log(handleOrder());

Line 1 returns tea. Line 2 calls it. Line 6 prints tea. After optimization, DevTools may reconstruct a serve frame from source positions, but its timing can shift because the physical work was merged.

Do not treat a reconstructed inlined frame as a lie. Treat it as a source-level explanation. Compare profiles before and after a code change, look at the surrounding path, and avoid assuming every visible frame is one machine-level call.

Here is a useful mental model. Before optimization, the machine can enter serve and return to handleOrder. After inlining, the instructions for returning tea can sit inside the caller, so the measured time may move up to the caller.

DevTools still has source positions and optimization metadata when available. It can display an inlined serve frame to help you reason in JavaScript names. Treat that display as a helpful reconstruction, not a promise about one physical stack frame.

Runtime call stats

Runtime call stats are V8 counters that group time spent in engine work. In Chrome tracing tools they can expose categories such as parsing, compiling, garbage collection, and inline-cache misses.

A runtime-call-stats reminderPop out in the code editor (opens in a new tab)JavaScript
// Chrome tracing can expose V8 runtime call stats.
// They group engine work such as parsing, compiling, GC, and IC misses.
// Use them to ask where engine time went, not to time one tiny call.
console.log("record a trace, then inspect RuntimeCallStats");

Line 1 is a comment, not a browser API call. Lines 2 and 3 name the kind of engine work these counters group. Line 4 prints a reminder: record a trace and inspect the RuntimeCallStats view.

Use this when application frames alone do not explain the cost. It is a diagnostic category view, not a promise that every engine version exposes identical counters or names.

For example, a profile can show a short application function that triggers a lot of parsing or garbage collection. Runtime call stats move the question one level lower: is the engine spending time compiling new code, collecting temporary objects, or handling inline-cache misses?

Keep this tool brief and targeted. First find the slow user action and its JavaScript path. Then inspect tracing categories only when the application frames leave a meaningful part of the cost unexplained.

Profiling a real problem

Start with a real action: typing in search, opening a cart, or loading a report. Reproduce it with realistic data. Record one baseline profile, locate a repeatable hot path, then change one idea and record again.

  • Profile a representative run, not an empty page.
  • Use a flame chart to find context and bottom-up self time to find busy leaves.
  • Keep the result of work alive so the optimizer cannot remove a benchmark.
  • Compare before and after profiles under similar conditions.
A slow-page starting pointPop out in the code editor (opens in a new tab)JavaScript
function filterOrders(orders, query) {
  return orders.filter((order) => order.name.includes(query));
}

const orders = [{ name: "tea" }, { name: "toast" }, { name: "tea cake" }];
const visible = filterOrders(orders, "tea");
console.log(visible.length);

Line 1 defines a filter used by a search page. Line 5 creates three visible orders. Line 6 filters for tea, and line 7 prints 2. The code is small on purpose: the same shape becomes expensive when a real page repeats it over a large list during every keystroke.

Find the hot function in a slow page

First, open the slow page with the normal amount of data. Start a recording, type one search term, and stop the recording after the results settle. Do not record an empty state if customers report slowness with a full list.

Second, use the flame chart to follow the wide path from the input event to a repeated function. Third, open bottom-up and look for the leaf with the largest self count. If that leaf is filterOrders, inspect how often it runs and how many orders it scans.

Finally, make one change, such as filtering only after a short pause or avoiding a duplicate filter. Record the identical action again. Keep the change only when the repeated hot path shrinks without breaking the result.

Sort the profiling ideas
  • Take one stack photo every few ticks
  • Record every function enter and exit
  • Merge repeated caller paths into bars
  • Sort leaf counts from largest to smallest
  • Choose a periodic observation interval
  • Add measurement around every call
Try it yourself
0 of 6 correct

Place each card by how it collects evidence or presents it.

Choose a category for every card. You can change an answer at any time; Reset clears them all.

For user-facing work, pair CPU evidence with measuring performance and main-thread. For compiler behavior, return to optimizing compilers.

Common mistakes

A profile is strongest when it changes a question into a smaller question. “Why is the page slow?” becomes “Why does this leaf appear so often for this action?” That keeps guesses out of the investigation.

  • “The widest frame is always the bug.” It may be a caller whose child owns the self time.
  • “One profile is a benchmark.” It is one observation of one workload.
  • “Inlining removes source context.” Tools can reconstruct source frames, though costs can move.
  • “A sample count is exact time.” It is an estimate shaped by the interval and run.
Profile words to keep separate
TermWhat it saysWhat it does not say
A wide barMany samples include that frameA promise that the source line is slow
Self timeSamples where a function is the leafTime including every child call
An inlined frameA source-level frame reconstructed by toolsProof the machine executed a normal call
A short missing functionWork that may fall between sample ticksProof that it never ran

When two runs disagree, check the workload, the machine state, and the recording length before changing the code. A new profile is evidence to compare, not a score to win.

Practice exercises

Exercise 1 · Warm-upPredict the sample count

Run the small model and type the printed count.

Starter codePop out in the code editor (opens in a new tab)JavaScript
const stacks = [["handleOrder", "makeTea"], ["handleOrder", "makeTea"], ["handleOrder", "serve"]];
console.log(stacks.filter((stack) => stack.at(-1) === "makeTea").length);

Answer, then press Check. Spacing and letter case don’t matter.

    Exercise 2 · Warm-upName the hot leaf

    The report has five makeTea leaf samples and one serve leaf sample. Which function should you inspect first?

    Answer, then press Check. Spacing and letter case don’t matter.

      Exercise 3 · PracticeChoose the Node output

      Choose the flag that writes a DevTools-readable CPU profile file.

      Answer, then press Check. Spacing and letter case don’t matter.

        Exercise 4 · PracticePredict optimized output

        What does the inlining model print?

        Starter codePop out in the code editor (opens in a new tab)JavaScript
        function serve() { return "tea"; }
        function handleOrder() { return serve(); }
        console.log(handleOrder());

        Answer, then press Check. Spacing and letter case don’t matter.

          Exercise 5 · ChallengeApply profiling to a checkout

          A checkout page feels slow with a large cart. What should you do before changing its rendering code?

          Answer, then press Check. Spacing and letter case don’t matter.

            Exercise 6 · ChallengeChoose an engine category

            Name one kind of engine work that runtime call stats can group.

            Answer, then press Check. Spacing and letter case don’t matter.

              Check your understanding

              Use the same rule for every question: first identify what the tool observed, then identify what the view groups.

              CPU profiling quiz · 8 questionsScore: first tries count
              1. Question 1 of 8What is a sampling profiler doing?

                Choose an answer to see the explanation.

              2. Question 2 of 8What does this print?

                Read the code, then predictPop out in the code editor (opens in a new tab)JavaScript
                const stacks = [["handleOrder", "makeTea"], ["handleOrder", "makeTea"], ["handleOrder", "serve"]];
                console.log(stacks.filter((stack) => stack.at(-1) === "makeTea").length);

                Choose an answer to see the explanation.

              3. Question 3 of 8What does a flame-chart bar width represent?

                Choose an answer to see the explanation.

              4. Question 4 of 8Why inspect bottom-up self time?

                Choose an answer to see the explanation.

              5. Question 5 of 8Which Node flag writes a .cpuprofile file?

                Choose an answer to see the explanation.

              6. Question 6 of 8What may happen after serve is inlined into handleOrder?

                Choose an answer to see the explanation.

              7. Question 7 of 8What are runtime call stats useful for?

                Choose an answer to see the explanation.

              8. Question 8 of 8A function appears in every stack but is never the leaf. What does that show?

                Choose an answer to see the explanation.

              Key takeaways

              • Sampling takes periodic stack photos and estimates where execution was observed.
              • Instrumentation records exact events but adds overhead.
              • Flame charts show merged call context; bottom-up views sort busy leaves.
              • Use `--cpu-prof` for a DevTools-readable Node profile and Linux perf when native stacks help.
              • Inlining can change physical frames while tools preserve source context.
              • Runtime call stats separate engine work from application work.

              Remember the one-liner.
              Profile the real action, then follow the repeated path before optimizing.

              Coming next: Benchmarking without fooling yourself.

              CompleteFrontend Clear concepts. Working examples.