What happens when you range over invalid UTF-8, and what does string(65) produce?
Question 12HardGo 1.22 to 1.25
s := "a\xffb"
fmt.Println(len(s)) // 3
for i, r := range s {
fmt.Printf("%d:%U ", i, r)
}
// 0:U+0061 1:U+FFFD 2:U+0062
fmt.Println(string(rune(65))) // "A"
fmt.Println(strconv.Itoa(65)) // "65"
fmt.Println(string([]rune{0x4E16, 0x754C})) // "世界"
fmt.Println(utf8.ValidString(s)) // false
When range hits a byte that isn't valid UTF-8, it yields utf8.RuneError (U+FFFD, the replacement character) and moves forward by one byte. It never panics. Converting a string to []rune does the same thing.
string(x) where x is an integer turns the value into the UTF-8 encoding of that code point. It does not give you the decimal text. So string(65) is "A". Since Go 1.15, go vet flags string(int) as a probable mistake. If you really want the character, write string(rune(x)). For the number as text, use strconv.Itoa or fmt.Sprint. Invalid code points, such as values above 0x10FFFF or surrogate halves, become "�".
What the interviewer wants to hear: this bug shows up in real code, for example building an ID with "user" + string(id).
More on Language Fundamentals & Types
- Q10In a type switch, what is the type of the bound variable when a case lists several types, and how does
case nilbehave? - Q11How are strings represented in Go? What does this print?
- Q13What does converting between
stringand[]bytecost, and how do you build strings efficiently? - Q14When are the arguments to a deferred call evaluated? What does this print?
- Q15In what order do deferred calls run, and what is wrong with using
deferinside a loop? - Q16How can a deferred function change a function's return value? What do
f()andg()return?