Skip to content
Haiku OSDeep Dive Published Updated 8 min readViews unavailable

Haiku Kernel Condition Variables: Experimental Wait and Notify Semantics

Understand Haiku kernel condition-variable entries, published waits, mutex handoff, notification results, timeout flags, and private API stability limits.

Haiku’s kernel condition-variable implementation provides a wait/notify mechanism for kernel threads and interrupt-related events. It is a narrow kernel-internal facility, not a supported userland synchronization class. The current documentation explicitly marks the API experimental and subject to change, and the header is under headers/private/kernel. Kernel contributors can use it to understand existing synchronization paths; application developers should prefer the documented public synchronization APIs rather than include a private kernel header.

A condition variable is not the condition

A condition variable is a queue of waiters associated with an opaque object. It does not store the application predicate for you. The predicate might be “device data is ready,” “a thread has exited,” or “a kernel object is being removed.” Protected state remains the source of truth. A waiter must check that state, register to wait without losing a concurrent notification, sleep, then check the predicate again after waking.

The “check, then wait” sequence is the classic lost-wakeup hazard. If a producer changes the predicate and sends a one-shot notification after the consumer checks but before it is attached to the wait list, a consumer can sleep indefinitely unless both sides use a synchronization protocol that closes that gap. Haiku provides a Wait(mutex*, flags, timeout) convenience method that is documented to add an entry, atomically release the mutex, wait, and reacquire the mutex. The producer must coordinate predicate changes with the same lock discipline.

Always wait in a predicate loop, not an if branch. A wakeup says to inspect state again; it does not prove the resource is still available or that the event remains true. Another thread may consume the resource first, NotifyAll() may wake multiple contenders, or the wake may carry a timeout or interrupt status.

Published and anonymous variables

ConditionVariable::Init() initializes an anonymous condition variable associated with an object pointer. It can be used directly through the object but is not available to ConditionVariableEntry::Add(object) lookup and does not appear in the kernel debugger’s published-variable list. Publish(object, objectType) registers a condition variable under an opaque object key, allowing other drivers or kernel modules to locate it by that key. The docs give device-tree nodes and hot-plug notifications as examples.

Publication requires careful identity and lifetime management. The key is a pointer identity, not a string name or a value copy. Publish only while the associated object is valid, prevent address reuse from accidentally identifying a replacement object, and unpublish when the association ends. objectType is a debugging description; it is not an access-control mechanism or a globally unique name.

An entry can be added only to one ConditionVariable. For a published object, ConditionVariableEntry::Add() returns false if no variable is published and records B_ENTRY_NOT_FOUND for a later wait. Check the boolean result rather than assuming a lookup succeeded. For an anonymous variable, use its direct Add() or Wait() path instead of published-object lookup.

Entry-based and convenience waits

ConditionVariableEntry exposes a two-step Add() then Wait() interface. It is useful when the caller must perform lock handoff or another ordering step between enrollment and sleeping. An entry can be added to only one variable at a time. The docs say it can be safely deleted without waiting for the event and can be reused after it has been removed from a variable.

The direct Wait(object, flags, timeout) convenience overload adds itself to a published variable and immediately waits. This removes the caller’s opportunity to do work between registration and the block. Use it only when that single-step ordering matches the protected predicate protocol. An anonymous variable uses the instance Wait() methods instead.

The mutex overload performs an atomic unlock-and-wait operation and reacquires the mutex before returning. The recursive-lock overload preserves its recursion count according to the documentation. These are kernel-level contracts; do not substitute a userland lock or an unrelated mutex type. If a kernel path must use a different synchronization design, the two-step entry API can be combined with that design, but the ordering and memory visibility must be proven at the call site.

Conceptual kernel pseudocode:

mutex_lock(&state->lock);
while (!state->ready) {
    status_t status = state->changed.Wait(&state->lock,
        B_RELATIVE_TIMEOUT, timeout);
    if (status != B_OK) {
        // Re-check state and handle timeout/interruption explicitly.
        break;
    }
}
mutex_unlock(&state->lock);

This is a control-flow sketch, not a complete compilable driver: initialization, lock ownership, absolute versus relative timeout calculation, and cleanup are omitted. Confirm the mutex API and allowed execution context for the exact kernel subsystem. The central rule is that ready is checked again while holding the same lock after each return.

Notification results and event timing

NotifyOne(result) wakes one attached entry; NotifyAll(result) wakes all attached entries. The return value is a count of entries that receive the result, so zero means no entry was attached at that instant. The status supplied by the notifier is what a notified wait returns. Use a conventional success status for ordinary wakeup and a documented error status for an abort condition; do not overload positive values without checking the consumer’s status interpretation.

The documentation calls out an important race: if an entry has already been added but has not yet entered its blocking wait, notification is still delivered and the later Wait() does not block. This protects the Add-then-Wait handoff. It does not turn notification into a durable event queue. A notification that occurs before a waiter adds itself finds no entry; the return count can be zero and the waiter must rely on the protected predicate when it checks state.

The static NotifyOne(object, result) and NotifyAll(object, result) methods find a previously published variable by pointer key. Check the returned count and ensure the object is still in the published lifetime window. If zero waiters is a valid state, it is not an error by itself. If the caller expected a waiter, it can indicate a lifecycle or ordering problem worth tracing.

Timeouts, interrupts, and status handling

Wait() accepts flags for relative or absolute timeout, and the documentation also lists flags that control interrupt behavior, including B_CAN_INTERRUPT and B_KILL_CAN_INTERRUPT. Relative timeouts are intervals; absolute timeouts are deadlines in the clock domain required by the kernel API. Do not pass a wall-clock timestamp where the API expects a monotonic deadline. Handle B_WOULD_BLOCK, interruption, cancellation, and notifier-supplied errors separately.

The entry implementation requires interrupts to be enabled when waiting and has a debug assertion that panics otherwise. Never call a blocking wait from an interrupt handler. An interrupt may notify an already-published variable, but the critical section and notification path must obey the kernel’s interrupt-safety constraints. Review the current implementation and call-site locks rather than assuming that every Notify method is safe from every interrupt level.

On timeout, the entry is removed from the wait list. If a notification races with timeout, the wait implementation gives precedence to a recorded notification status when one was actually received. Callers should still re-check their predicate and make state transitions idempotent; a timeout does not prove that the event cannot arrive immediately afterward.

Lifetime and teardown

The Doxygen contract states there are no restrictions on destruction order: a condition variable may be destroyed while entries wait on it, and entries may be destroyed before notification. That is an implementation guarantee for this private API, not permission to free the object associated with a published pointer while another subsystem may still call Notify(object, ...). Unpublish and coordinate the associated object’s lifetime before teardown.

The current implementation maintains a global lookup table for published variables and a per-variable entry list protected by kernel synchronization primitives. It tracks entry count and removes entries during wait completion or destruction. This is useful for understanding races but should not be copied into a driver as a substitute implementation. The data structures are private and can change.

For diagnostics, EntriesCount() reports attached entries, including entries that have been added but are still executing before their wait call. It is not identical to the number of threads currently blocked. Debug listing/dump methods help kernel investigation but are not an application monitoring interface.

Review and test checklist

For each use, identify the protected predicate, the lock or ordering protocol, publication lifetime, notification source, timeout clock, and status values. Test notify-before-add, notify-after-add-before-block, notify-while-blocked, timeout races, NotifyAll() with multiple waiters, object teardown, interruption, and no-waiter notifications. Confirm every path loops back to the predicate and releases locks exactly once.

Instrument state changes and notification counts rather than logging from an interrupt-sensitive path without checking that logging is safe. Include the object type, pointer identity where appropriate, wait status, timeout mode, and entry count in debug evidence. Avoid treating a zero notification count as proof of a lost event until the predicate and enrollment order are inspected.

Haiku’s condition variables are useful to understand kernel wait/notify behavior, but their status is explicit: private and experimental. Use the current source to review an in-tree kernel call site, keep predicate state separate from notification, use atomic lock handoff, and never expose this unstable implementation as a supported userland API.

Related:

Sources:

Comments