Haiku BEntryList: Directory Iteration, Cursor State, and Mutation Races
Traverse Haiku directories with BEntryList safely, preserve iterator state, handle symlinks deliberately, and avoid stale counts or snapshot assumptions.
BEntryList is the common iterator interface implemented by BDirectory and BQuery. It exposes entries in several representations, but those methods do not maintain independent cursors. GetNextEntry(), GetNextRef(), GetNextDirents(), and CountEntries() operate on the same iteration state. That detail can produce subtle bugs when code mixes convenience counting with a traversal loop or hands one iterator to multiple consumers.
For an ordinary directory walk, choose one representation, own the BDirectory for the duration of that walk, and treat enumeration as a changing view rather than a filesystem snapshot. A directory can be modified while it is being read. If correctness requires reconciling concurrent changes, combine an initial scan with node monitoring and a reconciliation pass rather than claiming that one enumeration is atomic.
Select the representation that matches the work
GetNextEntry() gives a BEntry object and has a traverse flag for deciding whether to follow a symbolic link. GetNextRef() returns an entry_ref, which is a compact reference to an entry within a directory. GetNextDirents() fills a caller-provided buffer with dirent records and can reduce object construction for lower-level scans. These calls all advance the shared iterator.
BDirectory directory(path);
status_t status = directory.InitCheck();
if (status != B_OK)
return status;
BEntry entry;
while ((status = directory.GetNextEntry(&entry, false)) == B_OK) {
// Inspect this entry without following symbolic links.
}
if (status != B_ENTRY_NOT_FOUND)
return status;
The false traverse argument is explicit: the loop sees the directory entry itself rather than asking the API to walk through a symbolic link. That is a conservative default for recursive tools. If the application does follow links, it needs a policy for cycles and repeated targets, not just a boolean flag. Use a set of stable node identities or another bounded traversal rule appropriate to the filesystem and application.
The loop distinguishes end-of-list from an actual error. GetNextEntry() reports B_ENTRY_NOT_FOUND when there are no more entries; do not collapse every non-B_OK result into normal end-of-directory. Preserve the error for logging and recovery. Also validate InitCheck() before iteration so an invalid BDirectory does not look like an empty directory.
Treat every BEntryList object as one cursor
The iterator position is shared across the retrieval methods. Calling CountEntries() in the middle of a GetNextEntry() loop can move or otherwise interfere with the same enumeration state. It also cannot provide a stable count if another process adds, removes, or renames entries while the loop runs. Keep counting and enumeration separate, and do not use a pre-count as a guarantee that exactly that many objects will be returned.
// Use a separate directory object if two consumers need independent scans.
BDirectory first(path);
BDirectory second(path);
// Each object owns its own iteration state; do not mix a count call into
// the same object that a consumer is already traversing.
int32 estimate = first.CountEntries();
status_t status = second.InitCheck();
This distinction matters in layered code. A helper that calls CountEntries() for a progress bar can silently perturb the caller’s directory scan if both functions receive the same BDirectory. Either count with a separate object, avoid pre-counting, or redesign progress reporting around the number actually processed. For unbounded or remote-like storage operations, displaying a moving count is often more honest than promising a percentage from an unstable estimate.
Use one directory object per active traversal. If recursion opens child directories, each child should have its own BDirectory instance while the parent keeps its position. Do not call Rewind() on the parent to implement recursion or to “restart” a nested consumer; that rewinds the same cursor and invalidates the caller’s expectation. BQuery also implements BEntryList, but its fetch and query lifecycle differ from directory enumeration; use the query-specific documentation for those rules.
Follow symlinks as an explicit policy
The GetNextEntry() traverse argument controls whether a returned symbolic link is followed. For inventory, cleanup, and backup scanning, defaulting to false avoids unexpectedly crossing into a different subtree. A recursive walk that follows links can revisit the same node through several paths or form a loop. It must track visited identities and apply a maximum depth or work budget.
Do not confuse an entry’s name with the identity of its target. A path may be renamed or removed after it is enumerated. If the operation must act on the exact object observed, keep and validate an appropriate entry or node reference and handle failure when the filesystem changes. A prior BEntry check is not a lease: another process can replace the path before a later open or mutation. Perform the operation and check its returned status rather than assuming the earlier metadata remains true.
When recursion should include directories but not cross symbolic links, use a non-traversing entry retrieval and inspect the entry type before opening a child BDirectory. A tool that follows links should state whether it follows file links, directory links, or both. Avoid silently inheriting a traversal default from a helper library.
Use GetNextDirents() without assuming fixed-size records
GetNextDirents() can return one or more directory records in a caller-owned buffer. The records are variable length. Advance from one record to the next using each dirent’s d_reclen field; do not index by sizeof(dirent), because d_name has variable storage and padding is part of the record layout.
#include <dirent.h>
alignas(dirent) char buffer[8192];
int32 count = directory.GetNextDirents(
reinterpret_cast<dirent*>(buffer), sizeof(buffer), 64);
if (count < 0)
return static_cast<status_t>(count);
if (count == 0) {
// End of this entry list.
} else {
char* cursor = buffer;
for (int32 index = 0; index < count; index++) {
dirent* item = reinterpret_cast<dirent*>(cursor);
// Consume item->d_name before advancing.
cursor += item->d_reclen;
}
}
The example shows the record walk, not a universal serialization format. Use the API’s returned count and buffer bounds, validate record lengths if handling untrusted or corrupted data, and avoid retaining pointers into the temporary buffer. A dirent’s name storage belongs to that buffer. If the next stage needs an independent representation, copy the name or construct an entry_ref before reusing the buffer.
For most application code, GetNextEntry() or GetNextRef() is easier to audit. Choose GetNextDirents() when profiling shows the lower-level interface is useful and the record handling is tested. Do not assume asking for many records guarantees a particular batch size; loop until the API indicates end-of-list or an error.
Design for concurrent directory changes
Directory enumeration is not a snapshot transaction. A rename during traversal can cause an entry to be missed, encountered under an old name, or observed inconsistently depending on filesystem timing. Creating and deleting entries while scanning has similar race windows. A successful scan means only that the iterator returned those entries during that interval, not that no other entries existed or that the same paths still exist afterward.
For a file indexer or synchronization tool, a practical design is an initial scan followed by node-monitor events and periodic reconciliation. Monitor events can be coalesced or lost across failure and restart, so they should accelerate convergence rather than replace the authoritative scan. On startup after downtime, rebuild or reconcile. For a one-shot cleanup utility, collect candidate references, display a reviewable plan, then revalidate each target immediately before the destructive action.
Do not hold a filesystem iterator open while waiting on user input, remote I/O, or a long-running copy. Keep the scan phase short, copy the minimum data needed for later work, and handle missing entries as expected operational outcomes. A path can become stale between discovery and use.
Acceptance checks for a traversal
Test an empty directory, a directory with mixed files and subdirectories, inaccessible entries, symbolic links, nested paths, and names containing whitespace or non-ASCII characters. If using GetNextDirents(), include records that force multiple buffer boundaries and validate that d_reclen advances exactly through the returned byte area. Verify end-of-list and error handling separately.
Add a test that mutates the directory while scanning. The expected result should be convergence or a clear per-entry failure, not a promise that one pass produces a transactionally complete inventory. Test that a helper calling CountEntries() cannot move the caller’s cursor, and that each recursive child traversal uses an independent BDirectory object.
BEntryList is intentionally a shared abstraction over directory and query iteration. Its cursor is stateful, its results are not a snapshot, and its data representations have different ownership costs. Keep one iterator owner per scan, choose one retrieval method, make symlink policy explicit, and treat concurrent filesystem mutation as normal rather than exceptional.
Related:
- How to Create BFS Attributes and Indexes for Fast Haiku Queries
- Live Queries: Searching Haiku’s File System Like a Database
Sources: