Go

How do you benchmark concurrent code with b.RunParallel, and why does sync.RWMutex often fail to scale for read-heavy workloads?

Question 404HardGo 1.22 to 1.25

b.RunParallel(body) starts GOMAXPROCS × SetParallelism(p) goroutines (p defaults to 1) and splits the b.N iterations among them; each goroutine loops on pb.Next(). The reported ns/op is wall-clock time divided by total iterations, so with perfect scaling ns/op drops as you add CPUs. You cannot use b.Loop inside the parallel body; pb.Next plays that role.

func BenchmarkCacheGet(b *testing.B) {
    for _, impl := range []struct {
        name string
        c    Cache
    }{
        {"mutex", NewMutexCache()},
        {"rwmutex", NewRWMutexCache()},
        {"sharded", NewShardedCache(64)},
    } {
        keys := make([]string, 1024)
        for i := range keys {
            keys[i] = strconv.Itoa(i)
            impl.c.Set(keys[i], i)
        }
        b.Run(impl.name, func(b *testing.B) {
            b.ReportAllocs()
            b.RunParallel(func(pb *testing.PB) {
                i := 0 // per-goroutine state: no sharing
                for pb.Next() {
                    impl.c.Get(keys[i&1023])
                    i++
                }
            })
        })
    }
}
// go test -run='^
  

 -bench=CacheGet -cpu=1,2,4,8,16 -count=10 | benchstat -

Typical surprise: rwmutex barely beats (or loses to) mutex at high core counts. Every RLock/RUnlock does an atomic add on the shared readerCount field, so all readers write the same cache line and it bounces between cores. RWMutex only wins when the read-side critical section is long compared to that cache-line transfer. Options that do scale:

  • Sharding: N independently locked maps selected by key hash (and padded if they are small).
  • Copy-on-write snapshot via atomic.Pointer[map[K]V]: readers do one atomic load with no writes; writers copy and swap. Great for rarely changing config or routing tables.
  • sync.Map: tuned for keys written once and read many times, or disjoint key sets per goroutine (since Go 1.24 it is backed by a concurrent hash-trie, which reduces contention on mixed workloads).

Pitfalls: shared state inside the benchmark body (a common counter, a shared *rand.Rand) becomes the bottleneck you end up measuring; math/rand/v2 top-level functions are safe and scale because they use per-thread state. Keep setup outside RunParallel, and always compare several -cpu values; a single number tells you nothing about scalability.

More on Performance, Profiling & Testing

All 38 Performance, Profiling & Testing questions