What is false sharing and how would you detect and fix it in Go?
CPUs cache memory in 64-byte lines (128 on some ARM). If two cores write to different variables on the same cache line, the line ping-pongs between cores via the coherence protocol, and throughput collapses even though there's no logical sharing.
type counters struct {
hits atomic.Int64 // same cache line as misses
misses atomic.Int64
}
// Fixed: pad each hot field onto its own line.
type paddedCounter struct {
n atomic.Int64
_ [56]byte // 8 + 56 = 64 bytes
}
type shardedCounter struct {
shards [16]paddedCounter
}
func (c *shardedCounter) Add(shard int, d int64) {
c.shards[shard%len(c.shards)].n.Add(d)
}
func (c *shardedCounter) Load() int64 {
var total int64
for i := range c.shards {
total += c.shards[i].n.Load()
}
return total
}
Detection: a parallel benchmark (b.RunParallel) that scales badly with -cpu=1,2,4,8 even though there are no locks; hardware counters via perf c2c on Linux. The runtime itself pads per-P structures, and the standard library uses an internal cpu.CacheLinePad type.
Gotchas: padding costs memory, so only do it for a small number of very hot, concurrently written fields; true sharing (every goroutine incrementing one atomic) has the same symptom and is solved by sharding, not padding. Interviewer is looking for: hardware-level reasoning and measuring scalability, not just single-thread ns/op.
More on Performance, Profiling & Testing
- Q394How do GOGC and GOMEMLIMIT interact, and how would you tune the GC for a latency-sensitive service?
- Q395Why does preallocating slices and maps matter? What does this print?
- Q397How does struct field ordering affect memory usage and performance?
- Q398How can you convert between []byte and string without allocating? When is it safe?
- Q399How does the race detector work, what does it cost, and what are its limits?
- Q400How do you separate unit tests from slow integration tests, and what are golden files?