Haiku BString: UTF-8 Length, Indexing, and Mutation
Use Haiku BString safely by separating byte lengths from UTF-8 character offsets, managing borrowed buffers, and testing mutation boundaries.
BString is Haiku’s mutable string value type in the Support Kit. Its String() method exposes a NUL-terminated byte sequence, and most of its basic storage/indexing model is byte-oriented. It also provides UTF-8-aware operations such as CountChars() and CountBytes(). Correct code makes the unit explicit at each boundary: a byte count, a UTF-8 character offset, and a user-perceived grapheme are not interchangeable.
That distinction matters in interfaces, filenames, text protocols, and persistence. A string that looks like one visible symbol can contain multiple Unicode code points; one code point can use multiple UTF-8 bytes; and some displayed letters combine a base character with a mark. The Support Kit’s character helpers do not replace a full Unicode grapheme segmentation or normalization library.
Bytes versus UTF-8 character counts
Length() returns the string length in bytes. CountChars() returns the number of UTF-8 characters as understood by the API. CountBytes(fromCharOffset, charCount) converts a range expressed in character offsets into a byte count. Its documentation warns that it does not validate that the requested range is within bounds, so callers must check offsets and counts first.
#include <String.h>
BString word("caf\xC3\xA9"); // UTF-8 bytes for "café"
int32 byteLength = word.Length();
int32 characterCount = word.CountChars();
int32 bytesForPrefix = word.CountBytes(0, 3); // "caf"
For the valid UTF-8 literal shown, Length() is five bytes and CountChars() is four characters. The prefix count is three bytes. These values are intentionally different. Never pass a CountChars() result to an API that expects a byte length, or use a byte offset as though it were a character position. Audit names such as offset, length, and position in your own helper functions and add suffixes such as Bytes or Chars where ambiguity is likely.
CountChars() is not a grapheme-cluster counter. For example, a letter followed by a combining accent can be two code points while a user sees one composed glyph; emoji sequences may include multiple code points joined into one displayed symbol. Cursor movement, selection, truncation, and “one character” limits intended for users may require a Unicode segmentation policy beyond BString’s basic UTF-8 helpers.
Keep String() as a borrowed view
String() returns a NUL-terminated const char* owned by the BString. You must not modify or free the returned pointer. It becomes invalid when the owning BString is deleted. Treat it as a short-lived borrowed view and reacquire it after any operation that may mutate, reassign, or destroy the owner; do not cache it in another object without a separate lifetime guarantee.
BString path("/boot/home/config/settings");
const char* borrowed = path.String();
UseImmediately(borrowed);
path << "/updated";
// Reacquire path.String() here; do not assume `borrowed` still describes it.
The exact invalidation time can depend on storage sharing and mutation details; code should not rely on an internal buffer staying at the same address. If a callee needs to retain text beyond the call, pass a BString value or make a copy into storage whose lifetime the callee owns. If a C API requires a mutable buffer, use the dedicated LockBuffer()/UnlockBuffer() contract and check the current header for its required size and balancing rules rather than casting away const from String().
When passing a pointer and length pair to a C API, use Length() for bytes and include a terminating NUL only if that API requires it. A common error is to pass CountChars() as the byte length for a UTF-8 string; another is to pass Length() to a function whose parameter is a count of code points. Keep the destination buffer size separate from source length and check for integer overflow when adding a terminator to a caller-provided allocation.
Mutation and indexing boundaries
Append() and the stream-style << operators are convenient for building strings, but repeated concatenation in a hot loop may allocate or copy. When the final size is known, reserve capacity through the documented API or build in a bounded buffer and measure the path before optimizing. Avoid retaining String() pointers across any mutation. A reference to the BString object itself also must not be concurrently mutated without synchronization.
Many methods in the broader API have byte-oriented and character-oriented forms. Read each declaration’s parameter names and documentation instead of assuming every method interprets offsets in UTF-8 characters. Methods that accept an explicit length may mean bytes, and C functions such as strlen() always count bytes up to NUL, not visible characters. Mixing BString and libc operations is safe only when the encoding and units are known.
If input may contain malformed UTF-8, define what the application will do: reject it, replace invalid sequences, or preserve the bytes as opaque data. Do not use a character count as validation. A valid-looking count does not establish that the string meets a protocol’s normalization or security requirements. Validate and normalize at a well-defined boundary using a library designed for that requirement.
BString does not automatically guarantee canonical Unicode normalization. Two strings can render identically and compare as different byte sequences if one uses a precomposed character and the other uses a base plus combining mark. Decide whether byte equality, case-insensitive comparison, normalized Unicode equality, or locale-aware collation is required. Do not silently change normalization for identifiers or paths without matching the relevant filesystem or protocol semantics.
Range arithmetic and large inputs
The public length methods return int32, so code should not assume that arbitrary unbounded input can be represented. Enforce application-level maximum sizes before concatenating large content, parsing a file, or allocating a second copy. When calculating offset + count, check for negative values and overflow before calling APIs whose docs say they do not validate bounds.
Before extracting or replacing a range, validate both the unit and the endpoint. If the UI supplies a character offset but the downstream byte API needs a byte position, call the appropriate UTF-8 byte-count helper only after validating that the range lies within CountChars(). For repeated range operations, maintain a clear mapping between bytes and code points or recompute it from the current immutable snapshot; stale offsets become wrong as soon as preceding text changes.
Do not split UTF-8 by arbitrary byte boundaries. Truncating at a byte count can cut a multi-byte sequence in half. If the output limit is a transport byte limit, walk code-point boundaries and stop before the next code point would exceed the limit. If it is a visual/user-character limit, grapheme-aware segmentation may also be required. Test ASCII, multi-byte characters, combining marks, and emoji sequences separately.
Copying, sharing, and threading
BString supports value-style copying, and the implementation may share backing storage internally. That is an optimization, not a contract that lets two threads mutate the same BString object concurrently. Treat each mutable instance as owned by one thread at a time or guard access with a lock. When sending a string to a worker, pass a value snapshot and do not also mutate it through another alias unless the class contract for the target version explicitly guarantees the usage.
If a worker returns a modified string, return a new BString or an immutable result message and let the owner replace its current value. This makes request ordering visible and avoids a worker changing text while a view reads String() for drawing. The lifetime of a borrowed C string must always be bounded by a live owning BString, regardless of any internal reference count implementation.
Persistence and protocol conversion
When serializing text, specify the encoding at the file or protocol boundary. BString does not add a BOM, normalize line endings, or decide how invalid byte sequences should be represented. A UTF-8 application format should document whether it accepts only valid UTF-8 and how it handles Unicode normalization. Keep binary payloads out of string-only assumptions; use byte buffers with explicit lengths if NUL bytes or arbitrary binary data are allowed.
Use BString fields for Haiku APIs that expect them, but do not confuse a method accepting BString with an automatic conversion to the remote service’s preferred encoding. Convert and validate deliberately. For paths, preserve filesystem encoding semantics and avoid applying URL encoding or user-text normalization to a native path.
Test the units in code review and automation
Create tests that assert Length(), CountChars(), and CountBytes() for ASCII and multibyte UTF-8 examples. Include a combining sequence and a multi-code-point emoji to demonstrate that “characters” is not the same as grapheme clusters. Test the documented out-of-range warning by ensuring application code rejects invalid values before calling CountBytes().
For borrowed-pointer behavior, make wrappers that call a C API synchronously, then mutate or destroy the BString and verify that no pointer was retained. Use a sanitizer-enabled build where available to expose use-after-free mistakes. Test long strings near the application’s documented size cap and every conversion path to/from C APIs. Verify whether the receiving API expects a NUL terminator and whether its length is bytes.
The safe mental model is: BString is a mutable UTF-8-oriented byte string with selected character-aware helpers. Keep units explicit, treat raw pointers as borrowed, and use dedicated Unicode processing when the feature needs grapheme, normalization, or locale semantics.
Related:
- Haiku BStringList: Ordered Collections, Copy Semantics, and Joining
- Haiku BFont: Text Measurement, Baselines, and Spacing Modes
Sources: