How do you implement health checks (liveness vs readiness) for a Go service running in Kubernetes, and how do they interact with graceful shutdown?
Liveness asks "is this process stuck?". When it fails, Kubernetes restarts the container. It should be cheap and should not check dependencies. If the database goes down and every pod fails liveness, Kubernetes restarts all of them over and over.
Readiness asks "should this pod receive traffic?". When it fails, the pod is removed from the Service endpoints. It can check critical dependencies and warm-up state. A startup probe protects slow-starting apps from being killed by the liveness probe.
Shutdown sequence: on SIGTERM, fail readiness first. Endpoint removal propagates asynchronously to kube-proxy and load balancers, so keep serving for a few seconds. Then call srv.Shutdown to drain in-flight requests, and do all of this inside terminationGracePeriodSeconds. (A preStop sleep hook does the same waiting job.)
func main() {
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM, os.Interrupt)
defer stop()
var ready atomic.Bool
mux := http.NewServeMux()
mux.HandleFunc("GET /livez", func(w http.ResponseWriter, _ *http.Request) { w.WriteHeader(http.StatusOK) })
mux.HandleFunc("GET /readyz", func(w http.ResponseWriter, _ *http.Request) {
if !ready.Load() {
http.Error(w, "not ready", http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
})
srv := &http.Server{Addr: ":8080", Handler: mux, ReadHeaderTimeout: 5 * time.Second}
go func() {
if err := srv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Fatal(err)
}
}()
ready.Store(true) // after caches are warm and connections are open
<-ctx.Done()
ready.Store(false) // 1. stop receiving new traffic
time.Sleep(5 * time.Second) // 2. let endpoint removal propagate
sctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
if err := srv.Shutdown(sctx); err != nil { // 3. drain in-flight requests
log.Printf("forced shutdown: %v", err)
}
// 4. flush telemetry, close the DB pool, and so on
}
Gotchas:
Shutdowndoes not wait for hijacked connections (WebSockets). Usesrv.RegisterOnShutdownfor those.- Make sure your process is PID 1 or runs under a proper init (use the exec form of
ENTRYPOINT), or SIGTERM never reaches it. - For gRPC, use
grpc_health_v1andGracefulStop.
More on Observability, Debugging & Production Operations
- Q568Explain gRPC in Go: unary vs streaming RPCs, interceptors, deadlines, status codes and error details.
- Q569How do you keep protobuf messages backward and forward compatible? Why must you never reuse field numbers?
- Q571How do you build a minimal, secure Docker image for a Go service (multi-stage build, scratch/distroless, CGO_ENABLED=0, CA certs, time zone data)?
- Q572How do you manage configuration in Go (env vars, files, flags)? How do you validate it and reload it safely at run time?
- Q573How do you configure TLS correctly in Go (crypto/tls MinVersion, certificate reloading, mTLS)?
- Q574How do you implement WebSockets in Go? How do you handle concurrent writes, ping/pong keepalives and backpressure?