How do you benchmark concurrent code with b.RunParallel, and why does sync.RWMutex often fail to scale for read-heavy workloads?
Question 404HardGo 1.22 to 1.25
b.RunParallel(body) starts GOMAXPROCS × SetParallelism(p) goroutines (p defaults to 1) and splits the b.N iterations among them; each goroutine loops on pb.Next(). The reported ns/op is wall-clock time divided by total iterations, so with perfect scaling ns/op drops as you add CPUs. You cannot use b.Loop inside the parallel body; pb.Next plays that role.
func BenchmarkCacheGet(b *testing.B) {
for _, impl := range []struct {
name string
c Cache
}{
{"mutex", NewMutexCache()},
{"rwmutex", NewRWMutexCache()},
{"sharded", NewShardedCache(64)},
} {
keys := make([]string, 1024)
for i := range keys {
keys[i] = strconv.Itoa(i)
impl.c.Set(keys[i], i)
}
b.Run(impl.name, func(b *testing.B) {
b.ReportAllocs()
b.RunParallel(func(pb *testing.PB) {
i := 0 // per-goroutine state: no sharing
for pb.Next() {
impl.c.Get(keys[i&1023])
i++
}
})
})
}
}
// go test -run='^