What causes "fatal error: all goroutines are asleep - deadlock!" and when does Go fail to detect a deadlock?
Question 194MediumGo 1.22 to 1.25
The runtime's checkdead triggers when every goroutine is blocked on something that no other goroutine can unblock (channel ops, select{}, sync primitives), and there are no timers or I/O that could wake anything up. Classic causes:
func main() {
ch := make(chan int)
ch <- 1 // no receiver exists: deadlock
fmt.Println(<-ch)
}
Others include ranging over a channel nobody closes, a WaitGroup.Wait() whose counter never reaches zero, and locking a non-reentrant mutex twice.
When it is NOT detected (the senior-level part):
- A partial deadlock, where some goroutines are stuck while others still run (an HTTP server, a ticker). The stuck ones just leak silently.
- Any goroutine blocked in a network poll or a syscall, or a pending timer, stops the check from firing.
- Programs using cgo, or goroutines in
time.Sleep.
So in production services, deadlocks show up as goroutine leaks and hangs, not crashes. Diagnose them with pprof goroutine dumps (/debug/pprof/goroutine?debug=2), SIGQUIT, or go.uber.org/goleak in tests. Go 1.26 adds an experimental goroutine-leak profile built on GC reachability.
More on Channels & select
- Q192Why is a nil channel useful in a
select? Show an example. - Q193What does this print? (select with a closed channel and a nil channel)
- Q195What does this program do? (unbuffered send in main)
- Q196How does
for rangeover a channel work, and what are its pitfalls? - Q197What are directional channel types and why use them?
- Q198How do you implement a timeout on a channel operation? Compare
time.After,time.NewTimer, andcontext.