Haiku BCollator: Locale-Aware Sorting, Natural Numbers, and Search Keys
Sort Haiku text with BCollator strengths, locale-specific rules, numeric ordering, reusable sort keys, and safe tie-breaking for distinct records.
Bytewise string order is not the same as the order people expect in a localized file list, contact picker, or settings panel. Haiku’s BCollator compares strings using locale-sensitive collation rules, with configurable strength and optional numeric sorting. It is a comparison policy, not a replacement for stable identifiers, Unicode normalization policy, or a database’s uniqueness constraint.
Obtain the collator for the intended locale
The Locale Kit’s BLocale::Default() represents the current default locale. Its GetCollator() method copies the locale’s collator into a caller-provided BCollator. Prefer that route when the UI should follow the user’s locale. The no-argument BCollator constructor exists, but the API documentation specifically recommends obtaining the default collator through BLocale because the locale knows the current settings.
BCollator collator;
status_t status = BLocale::Default()->GetCollator(&collator);
if (status != B_OK)
return status;
status = collator.SetNumericSorting(true);
if (status != B_OK)
return status;
int order = collator.Compare("chapter 9", "chapter 12");
With numeric collation enabled, digit runs can be compared by numeric order so a string containing 9 can sort before one containing 12. That is a user-interface ordering behavior, not an arithmetic parser; do not use it to validate a numeric field or compare identifiers where byte-level order is part of a protocol.
The explicit-locale constructor is appropriate when an application is intentionally rendering content for a locale other than the current UI locale. Keep that decision visible in the code. Sorting a user’s current UI in a hard-coded locale can produce an order that conflicts with the labels and conventions shown elsewhere.
Strength determines which strings compare equal
BCollator exposes primary, secondary, tertiary, and quaternary strengths (and an identical level in the header). Lower strengths can treat some differences as insignificant for ordering. The documentation gives accented letters and case as examples: a primary comparison can ignore accents, a secondary comparison can distinguish accents, and a tertiary comparison is case-sensitive. Choose the weakest strength that matches the user-facing task, not as a global default for every comparison.
This has an important consequence: Compare() returning zero does not prove two strings are byte-identical. If a list uses the collator as its only comparison key, distinct records can collapse into one equivalence class. A display sort should add a deterministic secondary key such as a stable record ID or original byte sequence when the collator returns equality. A uniqueness check should compare the actual canonical identifiers required by the application, not rely on collation equality.
SetStrength() returns a status and should be checked. If the collator could not be initialized, current implementation may fall back to bytewise strcmp() in Compare(). That fallback is not localized behavior. Surface initialization/configuration failure or use an explicitly documented fallback so a user does not assume that locale ordering succeeded.
Punctuation behavior needs a current-source caveat
The API exposes SetIgnorePunctuation() and IgnorePunctuation(). However, the current upstream Compare() and GetSortKey() implementations both contain TODO comments for applying the fIgnorePunctuation setting. Setting the flag is therefore not sufficient evidence that punctuation is ignored by current comparisons or sort keys. Do not promise that mode in a user interface or test only the getter. If punctuation-insensitive matching is a product requirement, implement and test a separate normalization/filtering policy or wait for an upstream implementation that supports it.
Even where a future implementation applies punctuation policy, ignoring punctuation can create collisions: re-enter and reenter may compare equally under a policy that removes the hyphen. Retain original strings for display and use a tie-breaker or stable identity for records. Never silently discard punctuation from user data just because a sort configuration ignores it.
Compare directly or build sort keys
Compare() is the straightforward choice when a sort routine compares a modest set of pairs. For repeated comparisons against many strings, GetSortKey() computes a modified key that the documentation says can be compared with strcmp() or a similar byte comparison. Check its status and cache the resulting key only while the locale and collator configuration remain unchanged.
Sort keys are derived values, not durable IDs. Rebuild them after the locale, collation strength, or numeric sorting policy changes. Do not persist a key as the sole representation of a user’s sort order across Haiku versions or ICU data updates. Store the original text and the sort policy, then derive the key at runtime.
Keep comparator configuration immutable during a sort. Changing numeric behavior or strength while a sorting algorithm is in progress can make comparisons inconsistent: the same pair may be ordered differently at two points in one operation. Configure a fresh collator, build a fresh set of keys, then swap the completed sorted model into the UI as one generation. If the user changes sort preferences mid-operation, cancel or discard the old result instead of merging two orderings.
When comparing two keys, use a bytewise comparator designed for the key representation, and do not decode the key as display text. Keep key storage separate from source strings. A clear record model can store {displayText, stableId, sortKey} and sort by sortKey, then by stableId when keys compare equal.
Threading and object lifetime
The BCollator API documentation warns that it is not thread-safe. Do not call one instance concurrently from multiple worker threads without synchronization. A simple design gives each sorting worker its own collator configured from the same locale policy; another design protects one shared instance with a lock. Avoid holding a global application lock across a long sort.
Also account for locale changes. A collator copied from BLocale::Default() is a snapshot of configuration at the time of the copy. If the application permits locale changes while running, obtain a refreshed collator and rebuild cached sort keys. Tag asynchronous sort requests with a generation and discard results produced for an older locale so a slower background sort cannot replace newer user-visible order.
Search is not automatically the same operation as sorting
Collation can help with user-facing ordering and certain equality checks, but substring search, prefix matching, case folding, diacritic-insensitive search, and filesystem matching are separate behaviors. Do not assume Compare() answers whether one string contains another. If you build accent-insensitive search, define how it relates to locale collation and preserve original text for the result display.
Sort order is also not an authorization or security boundary. Two visually or collationally similar names can still identify different records. Use stable internal IDs for actions and permission checks, and treat displayed strings as labels only.
Test the policy with realistic data
Build fixtures containing upper and lower case, accented letters, punctuation, composed/decomposed Unicode sequences, digit runs such as 2 and 10, empty strings, and duplicate display labels. Test the target locales the application supports rather than assuming English order. Verify comparator transitivity and stable tie-breaking, because a sort routine requires a consistent ordering relation.
Test fallback behavior if locale data initialization fails. Confirm no shared collator is concurrently accessed without a lock. Change locale and strength while background sorting is active and ensure stale results are discarded. For a list that stores user data, confirm that two collator-equal names remain separate records and that selection still targets the correct stable ID.
Add explicit regression cases for punctuation because the current implementation caveat is easy to miss: toggle the ignore flag, compare strings that differ only by punctuation, and compare their generated keys. If the results do not match the requested policy, keep punctuation-sensitive behavior or apply a separately tested transform. Capture the Haiku revision and locale in the test report so a future upstream fix can be detected without changing product expectations silently.
BCollator provides locale-aware comparison and reusable sort keys, but every policy choice has consequences. Use the user’s locale intentionally, select strength per task, make equality collisions deterministic, rebuild derived keys when settings change, and treat punctuation-ignore as unsupported until the current implementation actually honors it.
Related:
- Haiku’s Locale Kit: Catalogs, Formatting, and Runtime Language Selection
- Haiku BString: UTF-8 Length, Indexing, and Mutation
Sources: