Yes, indexing into a string is not available in Rust. The reason for this is that Rust strings are saved in a contiguous UTF-8 encoded buffer internally, so the concept of indexing itself would be ambiguous, and people would misuse it: byte indexing is fast, but almost always incorrect (when your text contains non-ASCII symbols, byte indexing may leave you inside a character / unicode code point, which is really bad if you need text processing), while code point indexing is not free because UTF-8 is a variable-length encoding, so you have to traverse the entire string buffer to find the required code point.
If you are certain that your strings contain ASCII characters only, you can use the as_bytes() method on &str which returns a byte slice, and then index into this slice:
let num_string = num.to_string();
// ...
let b: u8 = num_string.as_bytes()[i];
let c: char = b as char; // if you need to get the character as a unicode code point
If you do need to index code points, you have to use the chars() iterator:
num_string.chars().nth(i).unwrap()
As I said above, this would require traversing the entire iterator up to the ith code element.
Finally, in many cases of text processing, it is actually necessary to work with grapheme clusters rather than with code points or bytes. For example, many emojis are composed of multiple code points, but are perceived as one "character". With the help of the unicode-segmentation crate, you can index into grapheme clusters as well:
use unicode_segmentation::UnicodeSegmentation
let string: String = ...;
UnicodeSegmentation::graphemes(&string, true).nth(i).unwrap()
Naturally, grapheme cluster indexing into the contiguous UTF-8 buffer has the same requirement of traversing the entire string as indexing into code points.
Answer from Vladimir Matveev on Stack OverflowYes, indexing into a string is not available in Rust. The reason for this is that Rust strings are saved in a contiguous UTF-8 encoded buffer internally, so the concept of indexing itself would be ambiguous, and people would misuse it: byte indexing is fast, but almost always incorrect (when your text contains non-ASCII symbols, byte indexing may leave you inside a character / unicode code point, which is really bad if you need text processing), while code point indexing is not free because UTF-8 is a variable-length encoding, so you have to traverse the entire string buffer to find the required code point.
If you are certain that your strings contain ASCII characters only, you can use the as_bytes() method on &str which returns a byte slice, and then index into this slice:
let num_string = num.to_string();
// ...
let b: u8 = num_string.as_bytes()[i];
let c: char = b as char; // if you need to get the character as a unicode code point
If you do need to index code points, you have to use the chars() iterator:
num_string.chars().nth(i).unwrap()
As I said above, this would require traversing the entire iterator up to the ith code element.
Finally, in many cases of text processing, it is actually necessary to work with grapheme clusters rather than with code points or bytes. For example, many emojis are composed of multiple code points, but are perceived as one "character". With the help of the unicode-segmentation crate, you can index into grapheme clusters as well:
use unicode_segmentation::UnicodeSegmentation
let string: String = ...;
UnicodeSegmentation::graphemes(&string, true).nth(i).unwrap()
Naturally, grapheme cluster indexing into the contiguous UTF-8 buffer has the same requirement of traversing the entire string as indexing into code points.
The correct approach to doing this sort of thing in Rust is not indexing but iteration. The main problem here is that Rust's strings are encoded in UTF-8, a variable-length encoding for Unicode characters. Being variable in length, the memory position of the nth character can't determined without looking at the string. This also means that accessing the nth character has a runtime of O(n)!
In this special case, you can iterate over the bytes, because your string is known to only contain the characters 0–9 (iterating over the characters is the more general solution but is a little less efficient).
Here is some idiomatic code to achieve this (playground):
fn is_palindrome(num: u64) -> bool {
let num_string = num.to_string();
let half = num_string.len() / 2;
num_string.bytes().take(half).eq(num_string.bytes().rev().take(half))
}
We go through the bytes in the string both forwards (num_string.bytes().take(half)) and backwards (num_string.bytes().rev().take(half)) simultaneously; the .take(half) part is there to halve the amount of work done. We then simply compare one iterator to the other one to ensure at each step that the nth and nth last bytes are equivalent; if they are, it returns true; if not, false.
I understand there are three things:
-
bytes
-
scalar values
-
graphemes
And due to this, indexing into strings isn't as straightforward as other languages.
I have a large string with only ASCII characters. I do not wish to iterate over it. I just want to index into specific positions of interest. I don't need a range, I just need a single character.
So for example I want to achieve the following. (Python example)
s = "7GATTACA" #really long string n = int(s[0]) if n == 7: some_complex_operation(s[2],s[7], ..) #passing in more than just these two characters elif n == 4: some_complex_operation(s[2],s[4], ..) #passing in more than just these two characters
I am making this example up, so it's possible there is a workaround, but I'd urge you to play along for the sake of discussion.
I came across a few possibilities:
-
Create a string slice of length 1 for each character? (which can panic, but won't in my case)
-
User
char_indices()and usenth() -
Use
chars()andenumerate() -
Use
bytes() -
Use a crate? Let's say I don't want to use a crate for now (because my use case prevents it).
Correct me if I'm wrong, but in all of these options, I'm either creating extra allocations, or extra iterations over the string/iterator?
What is the best way to address this?
Thank you.
Implement Index<usize> for String and &str - libs - Rust Internals
Can slice but can't index an str
How to index a string vector
indexing - How can I find the index of a character in a string in Rust? - Stack Overflow
Although a little more convoluted than I would like, another solution is to use the Chars iterator and its position() function:
"Program".chars().position(|c| c == 'g').unwrap()
find used in the accepted solution returns the byte offset and not necessarily the index of the character. It works fine with basic ASCII strings, such as the one in the question, and while it will return a value when used with multi-byte Unicode strings, treating the resulting value as a character index will cause problems.
This works:
let my_string = "Program";
let g_index = my_string.find("g"); // 3
let g: String = my_string.chars().skip(g_index).take(1).collect();
assert_eq!("g", g); // g is "g"
This does not work:
let my_string = "プログラマーズ";
let g_index = my_string.find("グ"); // 6
let g: String = my_string.chars().skip(g_index).take(1).collect();
assert_eq!("グ", g); // g is "ズ"
You are looking for the find method for String. To find the index of 'g' in "Program" you can do
"Program".find('g')
Docs on find.
(update at the end)
[Edit: I'm certain I could get away with just using as_bytes, but I'm also taking the opportunity familiarize myself with the Unicode issues, since I've never really worked with that and it seems like Rust supports it well].
I'm a very experienced SW Engineer, but I've never had to work with Unicode stuff. I'm using last year's Advent of Code as an excuse to learn Rust.
I'm using some code from "Rust By Example" to read lines from a file. I'm pretty sure I understand this part; I can print the lines that are read in:
fn read_lines<P>(filename: P) -> io::Result<io::Lines<io::BufReader<File>>>
where P: AsRef<Path>, {
let file = File::open(filename)?;
Ok(io::BufReader::new(file).lines())
} My code is
if let Ok(lines) = read_lines(fname) {
for line in lines.map_while(Result::ok) {
// do stuff
}
}
I'm pretty sure that line is a std::String in my loop; if I'm wrong, please let me know. If a line of input is L34, how can I safely get the L and 34 as separate values? Most of what I see online talk about using chars() and the iterator, but that feels like getting the 34 would be very cumbersome.
Note: I'm avoiding using the word "character" since that seems to be ambiguous (byte vs grapheme).
Updated:
After the helpful responses below, and some looking, I realized that I needed to know string iterators better (I tried to think of them more like C++ iterators). I ended up with this:
if let Ok(lines) = read_lines(fname) {
for line in lines.map_while(Result::ok) {
let mut chars = line.chars();
let direction = chars.next().unwrap();
let num = chars.as_str();
println!("line: {} => {} + {}", line, direction, num);
}
}