When would you use go tool trace instead of pprof? What can it show that profiles cannot?
Question 374HardGo 1.22 to 1.25
pprof aggregates samples into "where is time spent"; the execution tracer records a timeline of events: goroutine create/start/block/unblock, syscalls, network poller wakeups, GC phases, STW pauses, heap size, and per-P scheduling. Use it for latency and concurrency questions that aggregates hide:
- Poor parallelism — Ps idle while work is queued, or everything serialized on one goroutine.
- GC assist / STW pauses causing tail latency.
- Scheduler latency: goroutines runnable but not running.
- Why one specific request took 800ms (with user annotations).
import "runtime/trace"
func handle(ctx context.Context, req *Request) {
ctx, task := trace.NewTask(ctx, "handleRequest")
defer task.End()
trace.WithRegion(ctx, "decode", func() { decode(req) })
trace.Log(ctx, "userID", req.UserID)
}
// capture: curl -o trace.out 'http://127.0.0.1:6060/debug/pprof/trace?seconds=5'
// or: go test -trace=trace.out
// view: go tool trace trace.out
Since Go 1.21/1.22 the tracer was rewritten: overhead dropped to ~1–2% and traces are chunked so they scale. Go 1.25 added trace.FlightRecorder, which keeps the last few seconds in a ring buffer so you can snapshot a trace after detecting a slow request. The trace viewer also derives "goroutine analysis" and "synchronization blocking profile" views.
More on Performance, Profiling & Testing
- Q372How would you detect and diagnose a goroutine leak?
- Q373What do the block and mutex profiles measure, how do you enable them, and how do they differ?
- Q375What is testing.B.Loop (Go 1.24) and why is it preferred over the classic b.N loop?
- Q376What does this benchmark report, and why is it wrong?
- Q377How do you run benchmarks rigorously and compare two implementations? Explain -benchmem, ReportAllocs, and benchstat.
- Q378Write an idiomatic table-driven test with subtests. Why is this the preferred Go style?