Go

How do you monitor Go runtime health in production (runtime/metrics, goroutine counts, GC pause, heap), and what alerts would you set?

Question 576MediumGo 1.22 to 1.25

runtime/metrics (Go 1.16+) is the stable, cheap API for runtime statistics. Unlike runtime.ReadMemStats, it does not stop the world. The Prometheus Go collector can export all of these metrics with collectors.WithGoCollectorRuntimeMetrics(collectors.MetricsAll). Useful metrics include:

  • /sched/goroutines:goroutines
  • /memory/classes/heap/objects:bytes (the live heap)
  • /gc/heap/goal:bytes
  • /sched/pauses/total/gc:seconds (a histogram, Go 1.22+)
  • /cpu/classes/gc/total:cpu-seconds
  • /sched/latencies:seconds (time goroutines wait to be scheduled)
  • /sync/mutex/wait/total:seconds
func sample() {
	s := []metrics.Sample{
		{Name: "/sched/goroutines:goroutines"},
		{Name: "/memory/classes/heap/objects:bytes"},
		{Name: "/gc/heap/goal:bytes"},
		{Name: "/sched/latencies:seconds"},
	}
	metrics.Read(s)
	for _, m := range s {
		switch m.Value.Kind() {
		case metrics.KindUint64:
			fmt.Printf("%s = %d\n", m.Name, m.Value.Uint64())
		case metrics.KindFloat64Histogram:
			h := m.Value.Float64Histogram()
			fmt.Printf("%s: %d buckets\n", m.Name, len(h.Counts))
		case metrics.KindBad:
			fmt.Printf("%s unsupported by this Go version\n", m.Name)
		}
	}
}

Alerts that matter:

  • A goroutine count that grows steadily over hours means a leak. Alert on the trend, not a fixed number.
  • RSS or heap near the container limit or GOMEMLIMIT means an OOM is coming.
  • GC CPU fraction above about 25% means allocation pressure.
  • p99 scheduler latency in the milliseconds means CPU starvation or throttling. Also check cgroup cpu.throttled.
  • OS thread count rising means blocked syscalls.
  • Mutex wait time rising means contention.

Pair these with SLO alerts on latency and errors, because runtime metrics explain symptoms but should rarely page anyone on their own.

Gotchas:

  • Expose net/http/pprof only on an internal port, since it leaks information and can be expensive.
  • Run continuous profiling (for example Pyroscope or Parca) so you already have data when an alert fires.

More on Observability, Debugging & Production Operations

All 14 Observability, Debugging & Production Operations questions