How does the Go CPU profiler work, and what is the difference between "flat" and "cum" in pprof output?
The CPU profiler is a sampling profiler: the runtime asks the OS for a timer signal (SIGPROF on Unix) at 100 Hz by default, and on each tick records the stack of the goroutine running on that thread. Each sample represents ~10ms of CPU. It only measures on-CPU time — a goroutine blocked in I/O, channel ops, or time.Sleep shows up nowhere. For "why is my request slow but CPU is idle" use the block profile or execution trace.
- flat: time spent in the function's own code (the function is the leaf frame).
- cum: time spent in the function plus everything it calls (appears anywhere on the stack).
(pprof) top10 -cum
(pprof) list parseRecord // annotated source with per-line costs
(pprof) peek json.Unmarshal // callers and callees
(pprof) web // call graph; or run with -http=:8080 for flame graph
A function with high cum but low flat is an orchestrator — optimize its callees. High flat means the hot loop is inside it. Use -diff_base=old.pprof to compare two profiles. Gotcha: inlined functions still appear in stacks (pprof reconstructs them), but very short profiles give noisy, statistically weak results — profile for 30s+ under realistic load.
More on Performance, Profiling & Testing
- Q369How do you expose pprof in a production service, and what are the security and design gotchas?
- Q371Explain inuse_space vs alloc_space in a heap profile. Which do you use to find a memory leak vs GC pressure?
- Q372How would you detect and diagnose a goroutine leak?
- Q373What do the block and mutex profiles measure, how do you enable them, and how do they differ?
- Q374When would you use go tool trace instead of pprof? What can it show that profiles cannot?
- Q375What is testing.B.Loop (Go 1.24) and why is it preferred over the classic b.N loop?