Linux Restartable Sequences: Per-CPU Fast Paths and Abort-Safe Updates
Learn the Linux rseq ABI, per-thread registration, migration-safe critical sections, abort paths, and diagnostics for production hot paths.
Per-CPU data can reduce cache-line contention in hot paths, but a thread can migrate between CPUs at any scheduling point. Reading a CPU number and then indexing a per-CPU array is not sufficient: the thread may move after the read and update the wrong CPU’s slot. Pinning every thread is one possible policy, but it couples correctness and placement to affinity management.
Linux restartable sequences, commonly called rseq, provide a different contract. A thread registers a small user-space area with the kernel. A carefully delimited instruction sequence may then perform a short per-CPU operation. If a scheduling event or signal interrupts that sequence before its commit point, the kernel redirects execution to its abort path. The application retries or uses a slower fallback. Rseq does not pin the thread and does not make an arbitrary block of C code transactional; it exposes a narrow ABI for sequences whose machine-code shape and recovery path follow strict rules.
The problem rseq solves
Suppose a worker wants to increment a counter belonging to its current CPU. A call such as sched_getcpu() can return a useful snapshot, but it does not reserve that CPU. Migration can happen between the call and the store. A conventional atomic increment avoids a lost update, but many CPUs updating one shared cache line can still create contention. Sharding data by CPU helps only if the operation cannot commit to a stale shard.
Rseq couples CPU identification to a bounded user-space sequence. The sequence reads the per-thread ABI state, records which critical-section descriptor is active, checks that its expected CPU still matches, performs the update, and reaches a designated commit boundary. If the kernel detects an interruption in the protected interval, it changes the return instruction pointer to the sequence’s abort handler. The handler must not report success; it retries or takes a correct fallback path.
This is useful for carefully designed per-CPU counters, allocator metadata, and other small data structures where the application can tolerate a retry path. It is not a general-purpose replacement for mutexes, futexes, atomics, CPU affinity, or read-copy-update. Each of those solves a different synchronization problem.
Registration belongs to each thread
The rseq() system call registers an ABI data area for the calling thread. The structure contains fields such as the current CPU identifier and a pointer to the active restartable-sequence descriptor. The registration is thread-scoped: creating another thread does not make that thread’s application-defined rseq area valid automatically. Runtime libraries may register the area during thread initialization, so applications and libraries must coordinate rather than independently trying to claim a second registration.
On glibc systems, the C library exposes the per-thread rseq area through its documented integration. Application code should use that supported runtime interface or a maintained rseq library instead of inventing a second registration protocol. Registration failure needs a functional fallback. ENOSYS indicates that the syscall is unavailable; EINVAL can mean invalid size, alignment, or flags; and EBUSY can indicate that the current thread already has a registration. A sandbox can also reject a syscall, so a policy denial should not be misreported as proof that the kernel lacks rseq.
For an initial deployment check, trace a representative process and its threads:
strace -f -e trace=rseq -o /tmp/rseq.trace -- ./worker --self-test
grep -E 'rseq\(' /tmp/rseq.trace
A successful registration call is evidence that the traced thread registered an rseq area. It does not show that the application actually executes a restartable sequence, that a sequence is correct, or that it improves throughput. A missing trace line is also inconclusive if the runtime registered before tracing began, the syscall filter is unsupported, or the program uses a different libc or registration library. Tracing perturbs timing, so do not use a strace run as a performance measurement.
Treat the sequence as a small state machine
The ABI descriptor records the critical-section start, its post-commit address, and an abort address, together with an architecture-specific signature and flags. The user-space sequence publishes the descriptor before entering its protected instructions. It then verifies that the CPU identity captured before the critical section still matches the ABI’s current value. The final data update is the commit point. An interruption before that point must take the abort path; a completed operation must not be replayed as though it had failed.
The following is deliberately pseudocode. It describes the control flow, not a compilable implementation. Real rseq sequences are commonly emitted with architecture-specific assembly or generated by a mature library because the kernel ABI validates exact instruction addresses and abort signatures.
retry:
expected_cpu = rseq.cpu_id_start
publish(active_critical_section_descriptor)
if expected_cpu != rseq.cpu_id:
clear_descriptor()
goto retry_or_fallback
perform_only_the_validated_per_cpu_update()
# The final instruction reaches the descriptor's post-commit address.
return success
abort_handler:
clear_or_reestablish_sequence_state()
goto retry_or_fallback
The kernel’s guarantee is bounded by that descriptor and the ABI rules. The sequence must not contain an unhandled branch out of the interval, call arbitrary functions, block, or perform an operation whose side effects cannot be safely retried. If the thread is interrupted after publishing the descriptor but before the operation commits, the abort handler must leave the data structure in a valid state. If a sequence has already reached its commit point, the application must treat the update as complete even if unrelated work is later interrupted.
Do not confuse CPU-local safety with global synchronization
Rseq addresses migration and interruption of one thread’s short sequence. It does not automatically serialize two different CPUs updating a shared object, publish data to readers on other CPUs, or establish arbitrary acquire/release ordering. A per-CPU slot is useful precisely because one CPU runs only one task at a time, but readers that aggregate slots still need a design for concurrent observation, object lifetime, memory ordering, and CPU hotplug.
If a counter is read while writers update it, decide whether the result may be approximate, whether a retryable snapshot protocol is needed, or whether stronger synchronization is required. If a pointer is published to another thread, use the relevant atomic ordering and lifetime mechanism independently of rseq. The rseq ABI does not turn relaxed loads and stores into a process-wide transaction.
CPU IDs are also not durable identities. CPU hotplug, affinity changes, namespace views, and runtime configuration can change which CPUs are available. Use rseq’s values only according to the documented ABI and check the required fields before accessing them. Never retain a CPU ID as a permanent array index across operations without a lifecycle and bounds policy.
Compatibility and the extended ABI
The rseq interface is optional at runtime. A program must be able to operate when the syscall is absent, blocked, or unavailable through its C library. It must also respect the size and alignment communicated by the runtime or kernel rather than assuming that the newest structure layout exists everywhere.
Current kernel documentation describes an optimized V2 registration mode on architectures that support the required generic entry code. That mode requires registering the feature size advertised to user space; a legacy registration size preserves older behavior and does not enable every extension. Treat AT_RSEQ_FEATURE_SIZE, AT_RSEQ_ALIGN, and runtime-provided structure information as ABI inputs, not as constants to copy from a header found on a build machine. Do not write fields the kernel owns. Newer scheduler-related extensions are separately gated and should not be assumed just because basic rseq registration succeeds.
The kernel selftests are the right reference for the low-level assembly contract. They cover architecture-specific instruction sequences and abort behavior that a prose example cannot validate. If a project does not have a reason to maintain assembly across its supported architectures, use a mature library and keep the fallback path well tested.
Validate correctness before measuring speed
Build tests for every supported architecture and ABI, including the exact compiler and runtime versions used in production. Exercise successful completion, CPU migration, signal delivery, preemption, thread creation and exit, registration failure, and fallback behavior. Test under CPU pressure and affinity changes. Where the algorithm permits it, inject interruptions near the beginning, inside the update, and just before the commit boundary. Confirm that an aborted attempt is not counted as a successful update and that retries cannot livelock under the workload’s assumptions.
Use strace only to inspect registration and syscall policy. Use the kernel rseq selftests and application-specific invariants for correctness. Benchmark the rseq path against a realistic atomic or locked baseline with the same data layout, workload, affinity policy, and reader behavior. Record CPU model, kernel, libc, compiler, and thread placement; a synthetic single-thread test cannot establish a multi-socket production win.
If rseq appears to stop working after a libc, container, seccomp, or kernel change, separate the layers: confirm the syscall is permitted, confirm each participating thread has a valid registration, confirm the expected ABI size is supported, and inspect whether the application takes its fallback. A fallback is not a failure if it preserves correctness; silently treating an unregistered sequence as successful is.
Rseq is most valuable when the hot path is small, retryable, and demonstrably contended. Keep the architecture-specific sequence narrow, let the runtime own thread registration, make abort and fallback behavior explicit, and measure the complete workload before keeping the added complexity.
Related:
- Linux Futexes in Practice: Compare-and-Block, Wakeups, and Lost-Wake Prevention
- Linux NUMA Memory Placement: CPU Affinity, Memory Policy, and Locality
Sources: