What is profile-guided optimisation (PGO) in Go and what does it actually change?
PGO (GA in Go 1.21) feeds a CPU profile from production back into the compiler. If a file named default.pgo is in the main package's directory, go build uses it automatically (-pgo=auto is the default). You can also pass -pgo=path/to/profile, or -pgo=off.
# 1. collect a representative CPU profile from production
curl -o cpu.pprof 'http://localhost:6060/debug/pprof/profile?seconds=30'
# 2. commit it next to main.go
cp cpu.pprof ./cmd/server/default.pgo
# 3. build as usual: PGO is applied automatically
go build ./cmd/server
What the compiler does with it:
- Hot-call inlining: functions on hot call edges get a much larger inlining budget, so bigger functions are inlined there. Because inlining feeds escape analysis, this can also turn heap allocations into stack allocations.
- Devirtualisation: a hot interface method call, or indirect call through a function value, gets a guarded direct call to the most common concrete target, which can then be inlined.
Typical gains are 2-14% CPU. Build time goes up, especially the first build without a cache.
Gotchas: the profile should come from the real workload. A profile from an older version of the source still mostly works, because matching is by function name and line offset, and stale parts are ignored. A profile made from a PGO-built binary is fine too; Go's design expects that iterative loop. Only CPU profiles are used, not heap profiles.
What the interviewer is looking for: you know PGO is cheap to adopt, that its main mechanism is inlining plus devirtualisation, and that it pairs with the escape-analysis story.
More on Memory, GC & Runtime Internals
- Q365What are the runtime representations of Go's built-in types? What does this print on 64-bit?
- Q366What is bounds-check elimination and how do you help the compiler do it?
- Q368How does the runtime preempt goroutines, and what are GC safe points?