Go

How do you implement health checks (liveness vs readiness) for a Go service running in Kubernetes, and how do they interact with graceful shutdown?

Question 570HardGo 1.22 to 1.25

Liveness asks "is this process stuck?". When it fails, Kubernetes restarts the container. It should be cheap and should not check dependencies. If the database goes down and every pod fails liveness, Kubernetes restarts all of them over and over.

Readiness asks "should this pod receive traffic?". When it fails, the pod is removed from the Service endpoints. It can check critical dependencies and warm-up state. A startup probe protects slow-starting apps from being killed by the liveness probe.

Shutdown sequence: on SIGTERM, fail readiness first. Endpoint removal propagates asynchronously to kube-proxy and load balancers, so keep serving for a few seconds. Then call srv.Shutdown to drain in-flight requests, and do all of this inside terminationGracePeriodSeconds. (A preStop sleep hook does the same waiting job.)

func main() {
	ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM, os.Interrupt)
	defer stop()

	var ready atomic.Bool
	mux := http.NewServeMux()
	mux.HandleFunc("GET /livez", func(w http.ResponseWriter, _ *http.Request) { w.WriteHeader(http.StatusOK) })
	mux.HandleFunc("GET /readyz", func(w http.ResponseWriter, _ *http.Request) {
		if !ready.Load() {
			http.Error(w, "not ready", http.StatusServiceUnavailable)
			return
		}
		w.WriteHeader(http.StatusOK)
	})
	srv := &http.Server{Addr: ":8080", Handler: mux, ReadHeaderTimeout: 5 * time.Second}

	go func() {
		if err := srv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
			log.Fatal(err)
		}
	}()
	ready.Store(true) // after caches are warm and connections are open

	<-ctx.Done()
	ready.Store(false)          // 1. stop receiving new traffic
	time.Sleep(5 * time.Second) // 2. let endpoint removal propagate
	sctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
	defer cancel()
	if err := srv.Shutdown(sctx); err != nil { // 3. drain in-flight requests
		log.Printf("forced shutdown: %v", err)
	}
	// 4. flush telemetry, close the DB pool, and so on
}

Gotchas:

  • Shutdown does not wait for hijacked connections (WebSockets). Use srv.RegisterOnShutdown for those.
  • Make sure your process is PID 1 or runs under a proper init (use the exec form of ENTRYPOINT), or SIGTERM never reaches it.
  • For gRPC, use grpc_health_v1 and GracefulStop.

More on Observability, Debugging & Production Operations

All 14 Observability, Debugging & Production Operations questions