Skip to content
Haiku OSDeep Dive Published Updated 6 min readViews unavailable

Haiku BBlockCache: Fixed-Size Reuse, Fallback Allocation, and Ownership

Use Haiku BBlockCache for reusable fixed-size blocks while managing fallback allocations, matching frees, cache capacity, and mutex contention.

BBlockCache is a Support Kit helper for reusing a bounded number of same-sized memory blocks. The constructor takes the number of blocks to preallocate, a block size, and an allocation mode: B_OBJECT_CACHE uses new[] and delete[], while B_MALLOC_CACHE uses malloc() and free(). Get() obtains storage and Save() returns it for reuse or frees it when it cannot be cached.

The cache is more constrained than a general allocator. Only requests whose size matches the configured block size use the free list. A request with another size allocates a separate block, and returning a mismatched block frees it rather than retaining it. This makes the class useful for predictable object sizes, but ineffective if callers pass a changing size or assume every Get() comes from preallocated storage.

Choose one size and allocation family

Pick the fixed block size from the actual data structure that needs reuse. Include the full number of bytes the caller will access, and keep the same size value for Get() and Save(). Choose the allocation type once and pair it consistently. A block created with new[] must be returned to an object cache; a malloc() block must be returned to a malloc cache.

BBlockCache cache(16, sizeof(WorkItem), B_OBJECT_CACHE);
void* memory = cache.Get(sizeof(WorkItem));
if (memory == NULL)
    return B_NO_MEMORY;

WorkItem* item = new (memory) WorkItem();
Process(*item);
item->~WorkItem();
cache.Save(memory, sizeof(WorkItem));

This sketch uses placement construction because the cache returns raw bytes. Destruct non-trivial objects before returning their storage. BBlockCache knows how to allocate and free memory, but it does not know the object’s C++ lifetime, constructor, destructor, or invariants. A mismatch between allocated storage size and the size passed back to Save() can bypass the pool and release memory immediately.

If the object is trivially destructible, document that assumption. If construction throws or Process() returns an error, use a cleanup path that destroys any constructed object and returns or frees the block exactly once. Avoid exposing the raw pointer after Save() because the next Get() may hand the same memory to another task.

Understand capacity and allocation fallback

blockCount controls how many free blocks are initially allocated and the maximum number of free blocks retained. When the free list is empty, Get() allocates another block rather than blocking for a cache slot. When a returned block matches the configured size and there is room, Save() puts it back; otherwise it frees it. A burst can therefore temporarily exceed the preallocated count, after which excess returned blocks are released.

This behavior prevents the cache from acting as a hard capacity limit. If the application must not exceed a memory budget, enforce that budget above the cache and handle allocation failure explicitly. Use Get()’s null result as a normal failure path; do not assume preallocation guarantees that an allocation can never fail under system memory pressure.

The destructor frees blocks on the free list. Blocks checked out and not returned remain caller-owned according to the official docs, so deleting the cache does not reclaim them. Before destruction, stop producers of new requests, wait for every task using a block to finish, then return or free outstanding blocks through the correct allocator. A cache’s destructor is not a substitute for joining worker threads.

Account for thread safety and contention

The source documentation calls BBlockCache thread-safe. The current implementation uses a pthread_mutex_t around the free list. That protects concurrent cache operations, but mutex use can block. Do not treat the class as wait-free or suitable for a strict real-time callback merely because it preallocates blocks.

If multiple threads share a cache, test contention at peak request rates and check that the objects stored in returned blocks are separately synchronized. Thread-safe memory reuse does not make the payload thread-safe. A producer must not continue accessing a block after returning it while a consumer may receive it again.

Consider per-thread caches only if the application can manage their independent capacities and teardown. Duplicating caches increases memory consumption and complicates ownership. For a small bounded workload, one cache and a measured lock may be simpler; for a deadline-sensitive path, preassign fixed slots and avoid acquiring a mutex in the callback.

Handle debug heaps and unusual requests

The current implementation detects an active debug heap and disables caching, allocating and freeing directly instead. This makes debugging safer but means performance and allocation reuse behavior differ from normal builds. Do not compare debug and release allocation benchmarks as if the cache were functioning identically.

The implementation also ensures each internal free-block node can hold its list metadata. Very small requested blocks may therefore have a physical allocation larger than the requested usable size. Treat the configured size as the API’s matching size, not as the exact allocator overhead or total footprint. Verify the current implementation if a future compatibility target makes byte-level memory accounting important.

Save() has no status return, so an application cannot ask whether a block was cached or freed through that call. If cache hit rate matters, instrument the call site with counts of requests, matching sizes, outstanding blocks, and expected capacity. Do not inspect private class fields or infer cache behavior from the fact that a pointer value was reused; allocators can return the same address for other reasons.

Consider whether returning a block can occur after the cache has begun destruction. The API does not provide a shared ownership or shutdown barrier for concurrent destruction. Quiesce all callers before the cache object leaves scope. Use a small owner object or lifecycle state to prevent late callbacks from accessing a cache whose mutex and free list no longer exist.

Requests whose blockSize differs from the configured size bypass the free list. A common bug is using sizeof(T) in one call and a manually calculated size in another; centralize the constant and test that the cache hit path is actually exercised. Instrument allocations and reuse counts outside the critical path so you can distinguish cache effectiveness from fallback behavior.

Failure-oriented verification

Test a full cache, empty cache, allocation failure, mismatched size, extra blocks returned above capacity, debug-heap mode, non-trivial object destruction, repeated cache destruction, and shutdown with checked-out blocks. Verify every memory object is freed using the matching allocation family and no pointer is used after Save().

Test with several threads and run race detection or debug heap diagnostics where available. Measure contention separately from general heap latency. If the cache is used for protocol buffers or media payloads, ensure callers initialize every byte that is transmitted and do not leak stale contents from a previous user of the block.

Include a canary or generation field in application-level test objects if stale-pointer reuse would otherwise be hard to detect. After Save(), poison or clear the payload when appropriate, then verify the next owner initializes all fields before reading. Avoid clearing large buffers inside latency-critical sections unless the cost is measured; a separate initialization stage may be safer.

Acceptance criteria

Accept a BBlockCache integration when allocation size and family are stable, objects are constructed and destroyed separately from raw storage, callers return every checked-out block, memory-budget enforcement exists above the cache, and mutex/fallback behavior is accounted for. Verify normal and debug-heap paths.

BBlockCache can reduce repeated allocation for fixed-size workloads. It does not cap total process memory, own checked-out blocks, or guarantee nonblocking access.

Related:

Sources:

Comments