Go

Walk through how you would investigate a service whose p99 latency doubled after a deploy.

Question 402HardGo 1.22 to 1.25

A structured answer matters more than any one tool:

  1. Correlate metrics: CPU, RSS, goroutine count, GC cycles/sec and GC CPU fraction (runtime/metrics), request rate, downstream latency. Is it CPU-bound, GC-bound, lock-bound, or waiting on I/O?
  2. Diff CPU profiles old vs new build under similar load: go tool pprof -diff_base=old.pprof new.pprof. Look for new hot functions — often an accidental regexp.MustCompile in a handler, JSON reflection, or logging at debug level.
  3. Allocations: -sample_index=alloc_space diff. A new allocation in the hot path increases GC frequency, and GC assists add latency directly to request goroutines.
  4. Contention: enable the mutex and block profiles — a new global lock or a shared sync.Map/cache under write load shows up here, not in CPU.
  5. Execution trace for a few seconds (or trace.FlightRecorder triggered on a slow request): see scheduler latency, STW pauses, goroutines waiting on the network, poor parallelism.
  6. Reproduce in a benchmark, fix, and prove it with benchstat and a canary.
// Handy: tag profile samples by endpoint so pprof can filter them.
func withLabels(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        pprof.Do(r.Context(), pprof.Labels("route", r.Pattern), func(ctx context.Context) {
            next.ServeHTTP(w, r.WithContext(ctx))
        })
    })
}
// (runtime/pprof) then: go tool pprof -tagfocus=route=/api/search cpu.pprof

Interviewer is looking for: measure before changing anything, knowing which profile answers which question, and that on-CPU profiles alone can't explain wait time. Also check the boring causes: GOMAXPROCS vs container CPU quota (throttling), changed timeouts/retries, or a dependency upgrade.

More on Performance, Profiling & Testing

All 38 Performance, Profiling & Testing questions