Go

A CPU profile shows lots of time in runtime.findRunnable, stealWork, futex and usleep. What is going on and how do you fix it?

Question 181HardGo 1.22 to 1.25

That profile shows scheduler overhead, not your code. It usually means the program creates or wakes many tiny goroutines, so Ps keep running out of work, spin looking for more, and park and unpark threads.

Why it happens:

  • Each new or newly ready goroutine may call wakep(), which starts a spinning M if there is an idle P. Spinning Ms loop through findRunnable and stealWork, checking other Ps' queues.
  • When they find nothing, they park on a futex (notesleep). The next wakeup needs a futexwakeup syscall. With microsecond-sized tasks, this park and unpark cycle costs more than the work itself.
  • Channel ping-pong between goroutines on different Ps adds wakeups and cross-core cache traffic.

Fixes:

  • Batch work. Send slices of items through a channel instead of one item at a time, or give each worker a chunk.
  • Use a fixed pool of long-lived workers instead of a go statement per tiny task.
  • Check GOMAXPROCS. An oversized value, such as 64 in a 2-CPU container before Go 1.25, makes the problem much worse. Lowering it cuts spinning.
  • Confirm with go tool trace. Many very short goroutine slices and frequent proc start and stop events are the signature.

What the interviewer wants to hear: more goroutines is not free parallelism. Each unit of work has to be big enough to pay for scheduling it.

More on Goroutines & the Scheduler

All 35 Goroutines & the Scheduler questions