Go

What happens when you range over invalid UTF-8, and what does string(65) produce?

Question 12HardGo 1.22 to 1.25
s := "a\xffb"
fmt.Println(len(s)) // 3
for i, r := range s {
	fmt.Printf("%d:%U ", i, r)
}
// 0:U+0061 1:U+FFFD 2:U+0062

fmt.Println(string(rune(65)))  // "A"
fmt.Println(strconv.Itoa(65))  // "65"
fmt.Println(string([]rune{0x4E16, 0x754C})) // "世界"
fmt.Println(utf8.ValidString(s)) // false

When range hits a byte that isn't valid UTF-8, it yields utf8.RuneError (U+FFFD, the replacement character) and moves forward by one byte. It never panics. Converting a string to []rune does the same thing.

string(x) where x is an integer turns the value into the UTF-8 encoding of that code point. It does not give you the decimal text. So string(65) is "A". Since Go 1.15, go vet flags string(int) as a probable mistake. If you really want the character, write string(rune(x)). For the number as text, use strconv.Itoa or fmt.Sprint. Invalid code points, such as values above 0x10FFFF or surrogate halves, become "�".

What the interviewer wants to hear: this bug shows up in real code, for example building an ID with "user" + string(id).

More on Language Fundamentals & Types

All 39 Language Fundamentals & Types questions