Data.String
String operations, byte-oriented and ASCII-accurate.
Built over the primitive UTF-8 string operations. Base includes this module.
Two indexing families exist and they do not cost the same. char_at, substring, and str_len are counted in codepoints, so each call walks the string’s UTF-8 encoding from the start and costs time proportional to the length. A scanner that advances an index one position at a time through them pays that walk on every step and is quadratic in its input, which is invisible on a line and ruinous on a file. byte_at and byte_len are constant-time raw-byte access and are the first choice for a scanner, a tokenizer, or a hash. Reach for the codepoint family only when the answer itself must be counted in characters.
Functions and Values
str_join
str_join : (String, List(String)) -> String
Join a list of strings, placing sep between adjacent elements.
str_join(", ", ["a", "b", "c"])
a, b, c
str_repeat
str_repeat : (String, Int) -> String
s repeated n times (the empty string when n <= 0).
str_repeat("ab", 3)
ababab
pad_left
pad_left : (String, Int) -> String
Right-align s to width w by prepending spaces (unchanged if already wider).
pad_left("42", 5)
42
pad_right
pad_right : (String, Int) -> String
Left-align s to width w by appending spaces (unchanged if already wider).
pad_right("42", 5)
42
lines_of
lines_of : (List(String)) -> String
Join a list of strings with newlines between them.
lines_of(["one", "two"])
one
two
occurs_at
occurs_at : (String, String, Int, Int) -> Bool
Helper for the substring queries: whether needle sits at byte offset j in s, comparing from position k.
starts_with
starts_with : (String, String) -> Bool
True when s begins with prefix.
starts_with("foo", "foobar")
true
ends_with
ends_with : (String, String) -> Bool
True when s ends with suffix.
ends_with("bar", "foobar")
true
index_of_go
index_of_go : (String, String, Int) -> Int
Helper for index_of: search for needle in s from byte offset j.
index_of
index_of : (String, String) -> Int
The byte offset of the first occurrence of needle in s, or -1 if absent.
index_of("bar", "foobar")
3
contains
contains : (String, String) -> Bool
True when needle occurs anywhere in s.
contains("oob", "foobar")
true
map_case
map_case : (String, Int, Buf, Bool) -> Buf
Helper for to_upper/to_lower: fold ASCII case mapping over s into a byte buffer (up selects upper- vs lower-casing).
to_upper
to_upper : (String) -> String
ASCII upper-case of s (non-letters unchanged).
to_upper("Hello")
HELLO
to_lower
to_lower : (String) -> String
ASCII lower-case of s (non-letters unchanged).
to_lower("Hello")
hello
ltrim_idx
ltrim_idx : (String, Int) -> Int
Helper for trim: the first non-whitespace byte index at or after i.
rtrim_idx
rtrim_idx : (String, Int) -> Int
Helper for trim: the index just past the last non-whitespace byte before i.
slice_bytes
slice_bytes : (String, Int, Int, Buf) -> Buf
Helper for trim: collect the bytes of s in [lo, hi) into buf.
str_slice
str_slice : (String, Int, Int) -> String
The bytes of s in [lo, hi), clamped to the string’s bounds.
The byte-indexed counterpart of substring: both endpoints are byte offsets and reaching one is constant time, so a scanner that slices as it advances stays linear where the codepoint form is quadratic. The endpoints must fall on character boundaries, which they do when they come from comparisons against ASCII bytes or from a span the compiler emitted; a window that splits a character is repaired rather than rejected, so the result is always a well-formed String.
The span shares the parent’s bytes rather than copying them, and the parent stays alive for as long as the span does, so slicing costs the same whether the string is three bytes or three megabytes.
str_slice("foobar", 3, 6)
bar
trim
trim : (String) -> String
Strip leading and trailing ASCII whitespace.
trim(" hi ")
hi
index_of_from
index_of_from : (Int, String, Int) -> Int
The index of character c in s at or after position i, or -1 if absent.
The codepoint index is the answer, so the walk that finds it is the walk this returns a position into: index_of is the byte-offset counterpart and is linear. Kept for the caller who needs a character position.
split_from
split_from : (Int, String, Int) -> List(String)
Helper for split: split s on c, starting from position i.
Splits at a character, so it addresses the string the way index_of_from answers. A caller splitting a large document wants byte offsets from index_of and slices from str_slice.
split
split : (Int, String) -> List(String)
Split s into the pieces between each occurrence of character c.
split(char_at(",", 0), "a,b,c")
[a, b, c]
str_of_char
str_of_char : (Char) -> String
The single-character string containing c.
str_of_char(chr(65))
A
chars_from
chars_from : (String, Int) -> List(Char)
Helper for chars: the characters of s from position i onward.
Decoding every character is what this is for, and each one is asked for by its character position; a byte-level pass would have to decode the encoding itself to answer the same question.
chars
chars : (String) -> List(Char)
The list of characters in s.
chars("hi")
[104, 105]