How do you monitor Go runtime health in production (runtime/metrics, goroutine counts, GC pause, heap), and what alerts would you set?
Question 576MediumGo 1.22 to 1.25
runtime/metrics (Go 1.16+) is the stable, cheap API for runtime statistics. Unlike runtime.ReadMemStats, it does not stop the world. The Prometheus Go collector can export all of these metrics with collectors.WithGoCollectorRuntimeMetrics(collectors.MetricsAll). Useful metrics include:
/sched/goroutines:goroutines/memory/classes/heap/objects:bytes(the live heap)/gc/heap/goal:bytes/sched/pauses/total/gc:seconds(a histogram, Go 1.22+)/cpu/classes/gc/total:cpu-seconds/sched/latencies:seconds(time goroutines wait to be scheduled)/sync/mutex/wait/total:seconds
func sample() {
s := []metrics.Sample{
{Name: "/sched/goroutines:goroutines"},
{Name: "/memory/classes/heap/objects:bytes"},
{Name: "/gc/heap/goal:bytes"},
{Name: "/sched/latencies:seconds"},
}
metrics.Read(s)
for _, m := range s {
switch m.Value.Kind() {
case metrics.KindUint64:
fmt.Printf("%s = %d\n", m.Name, m.Value.Uint64())
case metrics.KindFloat64Histogram:
h := m.Value.Float64Histogram()
fmt.Printf("%s: %d buckets\n", m.Name, len(h.Counts))
case metrics.KindBad:
fmt.Printf("%s unsupported by this Go version\n", m.Name)
}
}
}
Alerts that matter:
- A goroutine count that grows steadily over hours means a leak. Alert on the trend, not a fixed number.
- RSS or heap near the container limit or
GOMEMLIMITmeans an OOM is coming. - GC CPU fraction above about 25% means allocation pressure.
- p99 scheduler latency in the milliseconds means CPU starvation or throttling. Also check cgroup
cpu.throttled. - OS thread count rising means blocked syscalls.
- Mutex wait time rising means contention.
Pair these with SLO alerts on latency and errors, because runtime metrics explain symptoms but should rarely page anyone on their own.
Gotchas:
- Expose
net/http/pprofonly on an internal port, since it leaks information and can be expensive. - Run continuous profiling (for example Pyroscope or Parca) so you already have data when an alert fires.
More on Observability, Debugging & Production Operations
- Q574How do you implement WebSockets in Go? How do you handle concurrent writes, ping/pong keepalives and backpressure?
- Q575What do runtime/debug.SetMaxStack, SetMaxThreads, SetGCPercent and SetMemoryLimit do, and when would you use them?
- Q577How do you detect and fix high latency caused by lock contention, GC pressure, or connection-pool exhaustion using production telemetry?