Go

Implement a heartbeat/watchdog: a worker must signal liveness every T, or a supervisor restarts it.

Question 531HardGo 1.22 to 1.25

The supervisor starts the worker with a cancellable context and a beat callback. A timer is reset on every heartbeat. If the timer fires, the worker is considered hung: the supervisor cancels its context and starts a fresh one. The heartbeat uses a non-blocking send into a 1-buffered channel, so a slow supervisor can never stall the worker, and redundant beats are simply coalesced.

func Supervise(ctx context.Context, timeout time.Duration,
	work func(ctx context.Context, beat func()) error) {
	backoff := 100 * time.Millisecond
	for ctx.Err() == nil {
		wctx, cancel := context.WithCancel(ctx)
		beats := make(chan struct{}, 1)
		beat := func() {
			select {
			case beats <- struct{}{}:
			default: // a beat is already pending
			}
		}
		done := make(chan error, 1) // buffered: orphaned worker can still exit
		go func() { done <- work(wctx, beat) }()

		timer := time.NewTimer(timeout)
	watch:
		for {
			select {
			case <-beats:
				timer.Reset(timeout) // Go 1.23+: no drain needed
			case <-timer.C:
				log.Println("watchdog: no heartbeat, restarting worker")
				break watch
			case err := <-done:
				log.Println("worker exited:", err)
				break watch
			case <-ctx.Done():
				break watch
			}
		}
		timer.Stop()
		cancel() // ask the (possibly hung) worker to stop

		select { // back off before restarting
		case <-time.After(backoff):
		case <-ctx.Done():
		}
		backoff = min(backoff*2, 30*time.Second)
	}
}

// Worker side:
func worker(ctx context.Context, beat func()) error {
	t := time.NewTicker(time.Second)
	defer t.Stop()
	for {
		select {
		case <-ctx.Done():
			return ctx.Err()
		case <-t.C:
			beat()
			// ... do a unit of work ...
		}
	}
}

Discussion points:

  • A truly hung worker (a deadlock, or a syscall that ignores ctx) cannot be killed. Cancelling only orphans it, and the new instance may run alongside it. Decide whether you must wait for done before restarting (safer for exclusive resources) or restart immediately (better availability).
  • Beat from the work loop, not from a separate ticker goroutine. A heartbeat goroutine that runs independently keeps reporting "alive" while the real work is stuck.
  • Use exponential backoff to avoid crash loops, and reset it after a healthy period.
  • break watch is needed because a bare break inside select exits only the select.
  • Choose timeout to be a few multiples of the beat interval to tolerate GC pauses and scheduling jitter.

More on Classic Concurrency Coding Problems

All 16 Classic Concurrency Coding Problems questions