What does indexing and ranging over a UTF-8 string print?
s := "héllo"
fmt.Println(len(s)) // 6
fmt.Println(s[1]) // 195
fmt.Printf("%c %q\n", s[1], s[1:3]) // Ã "é"
for i, r := range s {
fmt.Print(i, ":", string(r), " ") // 0:h 1:é 3:l 4:l 5:o
}
fmt.Println()
fmt.Println(len([]rune(s)), utf8.RuneCountInString(s)) // 5 5
// s[0] = 'H' // compile error: cannot assign to s[0] (strings are immutable)
A string is an immutable sequence of bytes, usually UTF-8. len counts bytes. é (U+00E9) is two bytes, 0xC3 0xA9, so the length is 6. s[i] returns a byte: s[1] is 0xC3 = 195, and printed as a character it is U+00C3, Ã.
range over a string decodes runes. The index is the byte offset where each rune starts, which is why it jumps from 1 to 3. Invalid UTF-8 decodes to U+FFFD, one byte at a time.
Gotchas: s[:n] can cut a multi-byte character in half. Reversing a string byte by byte corrupts it; convert to []rune first, and even that breaks combining characters, for which you need grapheme segmentation. With n := 65, string(n) gives "A", not "65", and vet's stringintconv check flags it. Use strconv.Itoa(n), or string(rune(n)) if you really want the character.
More on Tricky Output & Code-Review Puzzles
- Q505What does this integer puzzle print?
- Q506Why do these two floating-point comparisons give different results?
- Q508This compiles. Why does it panic at runtime?
- Q509Why does errors.Is return different results here?
- Q510Go 1.23 range-over-func: what does this buggy iterator do?
- Q511What does this counter print, and how do you prove the bug?