Go

How do you detect and fix high latency caused by lock contention, GC pressure, or connection-pool exhaustion using production telemetry?

Question 577HardGo 1.22 to 1.25

Start from the symptom: p99 latency is up but CPU is not saturated, so time is being spent waiting. Traces tell you which span is slow. Then work out which kind of waiting it is.

Lock contention. Signs: /sync/mutex/wait/total:seconds rising, and the mutex and block profiles. Enable them with runtime.SetMutexProfileFraction(100) and runtime.SetBlockProfileRate(...), then run go tool pprof http://host/debug/pprof/mutex. Fixes:

  • Shrink critical sections, and never hold a lock during I/O.
  • Shard the lock, or use RWMutex for read-heavy data.
  • Use atomic.Pointer with copy-on-write.
  • Use per-P caches such as sync.Pool.

GC pressure. Signs: high GC CPU fraction, frequent cycles (GODEBUG=gctrace=1), and a latency tail that lines up with GC. The cause is usually mark assist, which makes goroutines that allocate do GC work, rather than stop-the-world pauses. Fixes:

  • Use the alloc profile (-sample_index=alloc_space) to find and remove the hot allocations.
  • Preallocate slices and reuse buffers.
  • Avoid pointer-heavy structures.
  • Raise GOGC, together with a GOMEMLIMIT safety net.

Pool exhaustion. Signs: requests queue up waiting for a connection. For databases, look at db.Stats(). For HTTP clients, remember that http.Transport keeps only 2 idle connections per host by default (MaxIdleConnsPerHost). Beyond that, connections churn and pile up in TIME_WAIT.

func exportPoolStats(ctx context.Context, db *sql.DB) {
	t := time.NewTicker(15 * time.Second)
	defer t.Stop()
	for {
		select {
		case <-ctx.Done():
			return
		case <-t.C:
			s := db.Stats()
			// WaitCount/WaitDuration rising = requests blocked waiting for a connection
			slog.Info("db pool", "open", s.OpenConnections, "in_use", s.InUse,
				"wait_count", s.WaitCount, "wait", s.WaitDuration, "max_open", s.MaxOpenConnections)
		}
	}
}

var client = &http.Client{
	Timeout: 5 * time.Second,
	Transport: &http.Transport{MaxIdleConns: 200, MaxIdleConnsPerHost: 50, IdleConnTimeout: 90 * time.Second},
}

Common causes of leaked connections are not closing resp.Body or rows, and transactions that are never committed or rolled back.

Execution traces (go tool trace) show all three kinds of waiting in one timeline: blocking, scheduler delay and GC. Go 1.25's trace.FlightRecorder keeps the last few seconds of trace in memory, so you can snapshot it at the moment a slow request is detected.

What the interviewer is looking for: you go from a hypothesis to the right profile or metric, you verify the fix, and you do not just throw more pods at the problem.

More on Observability, Debugging & Production Operations

All 14 Observability, Debugging & Production Operations questions