How do you instrument a Go service with Prometheus metrics (counters, gauges, histograms)? What are the label cardinality pitfalls?
Question 566MediumGo 1.22 to 1.25
Use the github.com/prometheus/client_golang library:
- Counters only go up (requests, errors). Query them with
rate(). - Gauges go up and down (in-flight requests, queue depth).
- Histograms put observations into buckets (latency, sizes). You compute quantiles on the server with
histogram_quantile, and you can aggregate across instances, which summaries cannot do.
var (
reqs = promauto.NewCounterVec(prometheus.CounterOpts{
Name: "http_requests_total", Help: "HTTP requests.",
}, []string{"route", "method", "code"})
inflight = promauto.NewGauge(prometheus.GaugeOpts{
Name: "http_in_flight_requests", Help: "In-flight requests.",
})
latency = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Help: "Request latency.",
Buckets: prometheus.DefBuckets,
}, []string{"route", "method"})
)
func instrument(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
inflight.Inc()
defer inflight.Dec()
rec := &statusRecorder{ResponseWriter: w, code: http.StatusOK}
start := time.Now()
next.ServeHTTP(rec, r)
route := r.Pattern // Go 1.23+: "GET /users/{id}", not "/users/42"
latency.WithLabelValues(route, r.Method).Observe(time.Since(start).Seconds())
reqs.WithLabelValues(route, r.Method, strconv.Itoa(rec.code)).Inc()
})
}
type statusRecorder struct {
http.ResponseWriter
code int
}
func (s *statusRecorder) WriteHeader(c int) { s.code = c; s.ResponseWriter.WriteHeader(c) }
// mux.Handle("GET /metrics", promhttp.Handler())
Cardinality pitfalls: Prometheus creates a separate time series for every unique combination of label values. Labels such as user ID, email, raw URL path, query string, request ID or error message create unbounded series. The result is memory exhaustion in Prometheus and in your own process, because the client keeps every series in memory forever.
- Use the route template, not the path.
- Bucket status codes (
2xx/5xx) if needed. - Keep labels to small, fixed sets.
- Use exemplars or traces for per-request detail.
Other gotchas:
- Name metrics with base units (
_seconds,_bytes) and give counters a_totalsuffix. - Choose histogram buckets around your SLO.
- Registering the same metric twice panics.
More on Observability, Debugging & Production Operations
- Q564How do you debug a Go program with Delve (dlv)? How do you attach to a running process, debug a test, and debug goroutines?
- Q565What does GOTRACEBACK control, and how do you get a core dump from a crashing Go program and analyze it?
- Q567How do you add distributed tracing with OpenTelemetry in Go? How does the trace context flow through context.Context and HTTP/gRPC?
- Q568Explain gRPC in Go: unary vs streaming RPCs, interceptors, deadlines, status codes and error details.
- Q569How do you keep protobuf messages backward and forward compatible? Why must you never reuse field numbers?
- Q570How do you implement health checks (liveness vs readiness) for a Go service running in Kubernetes, and how do they interact with graceful shutdown?