Go

What does this print? (len, indexing and ranging over a UTF-8 string)

Question 68MediumGo 1.22 to 1.25
s := "héllo"
fmt.Println(len(s), utf8.RuneCountInString(s))
fmt.Println(s[1], string(s[1]))
for i, r := range s {
	fmt.Printf("%d:%c ", i, r)
}
fmt.Println()
fmt.Println(string([]rune(s)[1]), s[1:3])

bad := "a\xffb"
for _, r := range bad {
	fmt.Printf("%U ", r)
}

Output:

6 5
195 Ã
0:h 1:é 3:l 4:l 5:o
é é
U+0061 U+FFFD U+0062

Explanation:

  • len counts bytes. The letter é (U+00E9) takes 2 bytes in UTF-8: 0xC3 0xA9.
  • Indexing s[1] gives a single byte (195). string(byte) treats 195 as a code point, which is U+00C3, "Ã".
  • range over a string decodes runes. The index is the byte offset, so it jumps from 1 to 3.
  • Invalid UTF-8 decodes to U+FFFD (the replacement character), one byte at a time.

Gotchas:

  • []rune(s) allocates 4 bytes per rune. Use it for random rune access, and use utf8.DecodeRuneInString for streaming.
  • Reversing a string by bytes corrupts multi-byte characters.
  • Even rune-level work splits grapheme clusters such as emoji with modifiers.
  • string(i) for an int variable i gives the character with that code point (i = 65 gives "A"), not the digits. go vet (stringintconv) flags it. Write string(rune(i)) if you mean a character, or strconv.Itoa(i) for digits.

More on Arrays, Slices, Maps & Strings

All 37 Arrays, Slices, Maps & Strings questions