What does this benchmark report, and why is it wrong?
Question 376HardGo 1.22 to 1.25
func square(x int) int { return x * x }
func BenchmarkSquare(b *testing.B) {
for i := 0; i < b.N; i++ {
square(42)
}
}
It typically reports something like 0.25 ns/op — roughly one clock cycle, i.e. an empty loop. square is trivially inlined, its argument is a constant, and the result is unused, so the compiler deletes the body entirely. You're benchmarking the loop counter.
Fixes, in order of preference:
// 1. Go 1.24+: b.Loop keeps call args/results alive.
func BenchmarkSquareLoop(b *testing.B) {
for b.Loop() {
square(42)
}
}
// 2. Pre-1.24: defeat constant folding and store to a package-level sink.
var sink int
func BenchmarkSquareSink(b *testing.B) {
x := 42
var r int
for i := 0; i < b.N; i++ {
r = square(x + i)
}
sink = r
}
Interviewer is looking for: suspicion of sub-nanosecond results, knowing the compiler can constant-fold and eliminate dead code, and checking with go test -gcflags=-m or by reading the assembly (go test -c then go tool objdump). Also: run with -count=10 and compare with benchstat rather than trusting one run.
More on Performance, Profiling & Testing
- Q374When would you use go tool trace instead of pprof? What can it show that profiles cannot?
- Q375What is testing.B.Loop (Go 1.24) and why is it preferred over the classic b.N loop?
- Q377How do you run benchmarks rigorously and compare two implementations? Explain -benchmem, ReportAllocs, and benchstat.
- Q378Write an idiomatic table-driven test with subtests. Why is this the preferred Go style?
- Q379What does this test print, and in what order?
- Q380What are the rules and pitfalls of t.Parallel()?