Explain 64-bit atomic alignment and false sharing. How do you lay out hot concurrent counters?
Atomic alignment: on 32-bit platforms (386, ARM, 32-bit MIPS), 64-bit atomic operations need 8-byte alignment, but int64 fields are only guaranteed 4-byte alignment. With the old atomic.AddInt64(&s.n, 1), the field had to come first in the struct or the program could panic. Since Go 1.19, atomic.Int64 / atomic.Uint64 are always 8-byte aligned on every platform, so use them.
False sharing: two independent variables in the same 64-byte cache line, written by different cores, cause the line to bounce between caches, and throughput collapses. Padding fixes it:
import (
"sync/atomic"
"golang.org/x/sys/cpu"
)
type ShardedCounter struct {
shards [8]struct {
n atomic.Int64
_ cpu.CacheLinePad // pads to cache line size (64 or 128 bytes)
}
}
func (c *ShardedCounter) Add(shard int, d int64) {
c.shards[shard&7].n.Add(d)
}
func (c *ShardedCounter) Load() (sum int64) {
for i := range c.shards {
sum += c.shards[i].n.Load()
}
return
}
What the interviewer is looking for: you can explain why a "lock-free" counter scales worse than expected, and you know to benchmark with -cpu=1,4,8. Also know the trade-off: padding costs memory, so use it only for contended hot fields.
More on Memory, GC & Runtime Internals
- Q347When does converting between
stringand[]byteNOT allocate? - Q348What does this print on a 64-bit platform, and why?
- Q350What are zero-sized types' memory semantics? What do these print?
- Q351What are the rules for valid
unsafe.Pointerusage? - Q352Why is keeping a pointer as
uintptrdangerous even if the object is still referenced elsewhere? - Q353How do you do zero-copy
[]byte↔stringconversion correctly, and what can go wrong?