Go

Why can many CPU-bound goroutines hurt performance, and how do you size a worker pool?

Question 173MediumGo 1.22 to 1.25

For CPU-bound work, only GOMAXPROCS goroutines can make progress at once. Adding more goroutines than that gives you no extra throughput. It does cost more:

  • Switching and preemption every 10 ms, which thrashes the cache.
  • More goroutines alive at once, so more memory held and more GC work.
  • Worse tail latency, because a given task waits behind many others.

Rules of thumb:

  • CPU-bound: use about runtime.GOMAXPROCS(0) workers. Do not use NumCPU, because it ignores container limits and user settings.
  • I/O-bound: use many more workers. Size them by the capacity of the downstream system (connection pool size, rate limits) and apply Little's law: concurrency ≈ throughput × latency.
  • Very small tasks: spawning a goroutine per 50 ns operation is pure overhead. Batch or chunk the input.
func parallelSum(xs []int) int {
	n := runtime.GOMAXPROCS(0)
	chunk := (len(xs) + n - 1) / n
	partial := make([]int, n)
	var wg sync.WaitGroup
	for w := range n {
		lo, hi := min(w*chunk, len(xs)), min((w+1)*chunk, len(xs))
		wg.Go(func() {
			for _, x := range xs[lo:hi] {
				partial[w] += x // false sharing risk; sum locally then store
			}
		})
	}
	wg.Wait()
	total := 0
	for _, p := range partial {
		total += p
	}
	return total
}

Follow-up: false sharing. Adjacent partial[w] entries sit on the same cache line, so the cores keep invalidating each other's caches. Sum into a local variable and write to the slice once, or pad the slice entries.

More on Goroutines & the Scheduler

All 35 Goroutines & the Scheduler questions