Implement a heartbeat/watchdog: a worker must signal liveness every T, or a supervisor restarts it.
Question 531HardGo 1.22 to 1.25
The supervisor starts the worker with a cancellable context and a beat callback. A timer is reset on every heartbeat. If the timer fires, the worker is considered hung: the supervisor cancels its context and starts a fresh one. The heartbeat uses a non-blocking send into a 1-buffered channel, so a slow supervisor can never stall the worker, and redundant beats are simply coalesced.
func Supervise(ctx context.Context, timeout time.Duration,
work func(ctx context.Context, beat func()) error) {
backoff := 100 * time.Millisecond
for ctx.Err() == nil {
wctx, cancel := context.WithCancel(ctx)
beats := make(chan struct{}, 1)
beat := func() {
select {
case beats <- struct{}{}:
default: // a beat is already pending
}
}
done := make(chan error, 1) // buffered: orphaned worker can still exit
go func() { done <- work(wctx, beat) }()
timer := time.NewTimer(timeout)
watch:
for {
select {
case <-beats:
timer.Reset(timeout) // Go 1.23+: no drain needed
case <-timer.C:
log.Println("watchdog: no heartbeat, restarting worker")
break watch
case err := <-done:
log.Println("worker exited:", err)
break watch
case <-ctx.Done():
break watch
}
}
timer.Stop()
cancel() // ask the (possibly hung) worker to stop
select { // back off before restarting
case <-time.After(backoff):
case <-ctx.Done():
}
backoff = min(backoff*2, 30*time.Second)
}
}
// Worker side:
func worker(ctx context.Context, beat func()) error {
t := time.NewTicker(time.Second)
defer t.Stop()
for {
select {
case <-ctx.Done():
return ctx.Err()
case <-t.C:
beat()
// ... do a unit of work ...
}
}
}
Discussion points:
- A truly hung worker (a deadlock, or a syscall that ignores ctx) cannot be killed. Cancelling only orphans it, and the new instance may run alongside it. Decide whether you must wait for
donebefore restarting (safer for exclusive resources) or restart immediately (better availability). - Beat from the work loop, not from a separate ticker goroutine. A heartbeat goroutine that runs independently keeps reporting "alive" while the real work is stuck.
- Use exponential backoff to avoid crash loops, and reset it after a healthy period.
break watchis needed because a barebreakinsideselectexits only the select.- Choose
timeoutto be a few multiples of the beat interval to tolerate GC pauses and scheduling jitter.
More on Classic Concurrency Coding Problems
- Q529Implement a readers-writer lock that prefers writers using only channels or sync.Mutex. Why is it hard to get right?
- Q530Implement a retry-with-timeout helper that runs a function in a goroutine and returns early when the per-attempt or overall deadline expires, without leaking goroutines.