How do you detect and fix high latency caused by lock contention, GC pressure, or connection-pool exhaustion using production telemetry?
Start from the symptom: p99 latency is up but CPU is not saturated, so time is being spent waiting. Traces tell you which span is slow. Then work out which kind of waiting it is.
Lock contention. Signs: /sync/mutex/wait/total:seconds rising, and the mutex and block profiles. Enable them with runtime.SetMutexProfileFraction(100) and runtime.SetBlockProfileRate(...), then run go tool pprof http://host/debug/pprof/mutex. Fixes:
- Shrink critical sections, and never hold a lock during I/O.
- Shard the lock, or use
RWMutexfor read-heavy data. - Use
atomic.Pointerwith copy-on-write. - Use per-P caches such as
sync.Pool.
GC pressure. Signs: high GC CPU fraction, frequent cycles (GODEBUG=gctrace=1), and a latency tail that lines up with GC. The cause is usually mark assist, which makes goroutines that allocate do GC work, rather than stop-the-world pauses. Fixes:
- Use the alloc profile (
-sample_index=alloc_space) to find and remove the hot allocations. - Preallocate slices and reuse buffers.
- Avoid pointer-heavy structures.
- Raise GOGC, together with a
GOMEMLIMITsafety net.
Pool exhaustion. Signs: requests queue up waiting for a connection. For databases, look at db.Stats(). For HTTP clients, remember that http.Transport keeps only 2 idle connections per host by default (MaxIdleConnsPerHost). Beyond that, connections churn and pile up in TIME_WAIT.
func exportPoolStats(ctx context.Context, db *sql.DB) {
t := time.NewTicker(15 * time.Second)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s := db.Stats()
// WaitCount/WaitDuration rising = requests blocked waiting for a connection
slog.Info("db pool", "open", s.OpenConnections, "in_use", s.InUse,
"wait_count", s.WaitCount, "wait", s.WaitDuration, "max_open", s.MaxOpenConnections)
}
}
}
var client = &http.Client{
Timeout: 5 * time.Second,
Transport: &http.Transport{MaxIdleConns: 200, MaxIdleConnsPerHost: 50, IdleConnTimeout: 90 * time.Second},
}
Common causes of leaked connections are not closing resp.Body or rows, and transactions that are never committed or rolled back.
Execution traces (go tool trace) show all three kinds of waiting in one timeline: blocking, scheduler delay and GC. Go 1.25's trace.FlightRecorder keeps the last few seconds of trace in memory, so you can snapshot it at the moment a slow request is detected.
What the interviewer is looking for: you go from a hypothesis to the right profile or metric, you verify the fix, and you do not just throw more pods at the problem.
More on Observability, Debugging & Production Operations
- Q575What do runtime/debug.SetMaxStack, SetMaxThreads, SetGCPercent and SetMemoryLimit do, and when would you use them?
- Q576How do you monitor Go runtime health in production (runtime/metrics, goroutine counts, GC pause, heap), and what alerts would you set?