Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux Futexes in Practice: Compare-and-Block, Wakeups, and Lost-Wake Prevention

Understand Linux futex wait and wake contracts, private versus shared keys, spurious returns, memory ordering, and safe synchronization tests.

A futex is not a kernel-owned mutex object. It is a Linux system-call mechanism built around a four-byte, aligned userspace word. In the uncontended path, a lock or condition variable can often be manipulated with ordinary atomic instructions and never enter the kernel. A thread calls futex() when it must block or wake another thread. This split makes futexes efficient, but it also means the application or runtime library owns the synchronization state machine.

The central primitive is compare-and-block: FUTEX_WAIT atomically checks that the userspace word still equals an expected value and only then queues the caller to sleep. If the value changed before the kernel committed the wait, the call returns instead of sleeping indefinitely. This closes a specific lost-wakeup race. It does not make a complete lock algorithm, guarantee fairness, or replace correct atomic memory ordering.

The word is userspace state

The futex word is a 32-bit value even in a 64-bit process, and it must be aligned on a four-byte boundary. Applications typically use an atomic integer type whose representation and alignment satisfy the ABI, then pass its address to the syscall wrapper. The kernel does not require a create or destroy call for an ordinary futex. Internal kernel state exists while waiters or other active operations need it and is keyed from the address and sharing mode.

This distinction explains why futexes are attractive for mutexes: the fast path can use a compare-and-exchange in userspace. But it also rules out casual use of a 64-bit pointer as the futex word or use of a word at an unaligned byte offset. Validate layout explicitly, especially when a word lives in a packed structure, a shared-memory ABI, or a language runtime’s custom object header.

The FUTEX_PRIVATE_FLAG tells the kernel that the synchronization word is used only between threads in one process. The private form allows process-local keying optimizations. It is incorrect for a word used by independent processes through shared memory, because those processes may map the same physical word at different virtual addresses. Conversely, omitting the private flag when the object is strictly process-local can cost performance but preserves the shared-capable interpretation.

Wait is conditional, not a notification queue

The canonical protocol is a loop around a predicate. A waiter first examines state with an atomic operation. If progress is impossible, it calls FUTEX_WAIT with the value that it expects to remain unchanged. A producer changes the state with an atomic operation and calls FUTEX_WAKE when it may have enabled progress. The wait can return because the value no longer matched, because a signal interrupted it, because a timeout expired, or because a wake selected it. The caller must reload and re-evaluate the actual predicate in every case.

while predicate_is_false():
    expected = atomic_load_acquire(state)
    if predicate_became_true(expected):
        break
    result = futex_wait_private(&state, expected)
    if result is interrupted_or_timed_out:
        handle_or_retry_according_to_policy()

This sketch is intentionally not a complete C lock implementation. A real algorithm must define its state transitions, compare-and-exchange ordering, cancellation semantics, timeouts, and wake cardinality. In particular, a wake is not an ownership transfer: waking one waiter does not reserve the lock for that thread. Another runnable thread may acquire the state first, and the woken thread must loop.

The futex wait operation makes its value check and transition to sleep atomic with respect to a concurrent state change and wake. That is the mechanism that prevents a thread from missing a producer transition between checking the word and entering the wait queue. The application’s release/acquire atomics still establish the visibility ordering for protected data. Do not assume that calling FUTEX_WAKE alone publishes a preceding data structure update with the language-level ordering your program needs.

Return values are part of the state machine

Treat EAGAIN from a wait as “the word no longer equals the expected value,” not as a fatal synchronization error. EINTR means a signal interrupted a blocking call. Timeouts, cancellation, and wakeups also need explicit policy. Even a return that appears successful only says that the wait ended; the predicate may already be false again by the time the caller runs. Every retry must recheck state under the required atomic ordering.

Wake counts must match the condition being signaled. Waking one waiter is appropriate when one unit of work becomes available and only one worker should claim it. Waking all may be necessary for a global state change such as shutdown, but creates a thundering herd if many threads race for a single resource. A semaphore-like counter, condition variable, or queue may express the state more clearly than a custom futex protocol.

Futex operations can be combined with flags that change keying or operation behavior. The realtime-clock modifier is accepted only by specific timed operations, and additional commands such as requeue and priority-inheritance operations have distinct contracts. Check the target kernel’s man page and headers before using one. Do not infer that every command is supported on every distribution just because the C header defines its number.

Condition variables and requeueing

Most applications should use pthread mutexes and condition variables rather than implementing futex operations directly. A condition variable is not a stored event: a signal is not remembered as a durable token if there are no waiters, which is why the predicate must be protected by a mutex and checked in a loop. POSIX libraries build that contract on internal mechanisms that can include futexes, but the internal implementation is not the application ABI.

Linux provides futex requeue operations that can move waiters from one futex address to another without waking every waiter to contend immediately. This is useful for some condition-variable designs, but requeueing is not a generic “move thread” operation. Priority-inheritance requeue variants have extra coupling to PI mutex state. If a library or runtime uses these operations, use its documented contract and stress the exact timeout, cancellation, and priority cases rather than reproducing a partial protocol from a syscall trace.

For process-shared synchronization, place the word in a shared mapping and use the non-private futex operation. Both processes must agree on the same state encoding and lifecycle. If one process crashes while holding a plain lock, a futex does not automatically repair the abandoned state. Robust pthread mutexes have a separate owner-death protocol; ordinary futex wait/wake alone does not provide recovery.

Trace without mistaking implementation detail for API

strace can reveal when a process enters the futex syscall, which operation it requests, and whether the call returns with an error. It cannot show the userspace atomic fast path that did not enter the kernel. The trace also reflects the particular libc, language runtime, and kernel combination, so it is evidence about that process rather than a stable promise that a condition variable always maps to one operation.

strace -ff -e trace=futex -o /tmp/futex.trace -- ./service --self-test
rg 'FUTEX_(WAIT|WAKE)' /tmp/futex.trace.*

Run this against a bounded test process, not a long-running production service with sensitive command-line arguments or data in its trace. Tracing changes timing and can hide or create contention. Pair the syscall trace with application-level counters for waits, wakeups, retries, timeouts, queue depth, and time spent blocked.

For an intermittent stall, record the futex address only as a process-local diagnostic token; addresses can be reused and are not stable object identifiers. Collect thread stacks, lock ownership metadata from the runtime if available, and the application predicate. A futex wait in a trace does not prove deadlock. The owner may be running, descheduled, blocked in I/O, or no longer able to release the state.

Failure modes to test

Test a waiter arriving just before a producer change, during the compare-and-block window, and after a wake. Verify that the state predicate prevents lost progress in all three cases. Exercise signals, timeout expiration, shutdown, cancellation, and object destruction while threads are waiting. Include more waiters than work items and verify that unused waiters sleep again rather than consuming nonexistent work.

Use ThreadSanitizer or another race detector for the userspace state machine where supported, then run stress tests on the actual Linux architecture and kernel family. A clean sanitizer run is not proof of kernel futex semantics, and emulated architectures may schedule races differently. Add fault injection for owner-thread exit and process-shared mappings if those are supported cases.

Production acceptance criteria

Write down the futex word’s state encoding, alignment, memory-ordering rules, process-sharing model, and wake policy. Define which return codes cause retry, cancellation, or failure. State who owns a lock at every transition and how an owner death is detected. Ensure the state object remains mapped for every waiter and wake call; unmapping or reusing storage while another thread still waits can create confusing failures even when the syscall itself behaves correctly.

Track contention rate and wait duration under the expected load. A futex can reduce CPU consumption during waits but cannot compensate for an overloaded critical section or unfair work queue. Compare the custom primitive against pthread synchronization using identical workload and correctness tests. Keep the primitive only if its measured advantage justifies the maintenance cost of a Linux-specific concurrency algorithm.

Related:

Sources:

Comments