Linux RCU in Practice: Read-Side Sections, Grace Periods, and Reclamation
Design Linux kernel RCU readers and updates safely by separating pointer publication from reclamation, selecting the right grace-period API, and testing lifetimes.
Read-copy-update (RCU) is a kernel synchronization technique for data structures that are read frequently and changed less often. Its central benefit is not that readers acquire a cheaper mutex. It is that readers can inspect a published object while an updater replaces it, provided the updater postpones freeing the old object until every reader that could have seen it has finished.
That is a lifetime protocol with three distinct responsibilities: protect a read-side traversal, publish a fully initialized replacement, and reclaim removed storage after a grace period. RCU does not serialize writers, make mutable fields race-free, or make a pointer safe after the read-side critical section ends. Most RCU bugs come from treating one of those separate rules as implicit.
The read-side contract
For normal RCU, readers bracket pointer lookup and every dereference of the protected object with rcu_read_lock() and rcu_read_unlock(). rcu_dereference() supplies the compiler and CPU ordering required to observe an object published through RCU. A reader must finish using that pointer before unlocking; copying the pointer into a long-lived work item and dereferencing it later defeats the protection.
static u32 active_mtu(void)
{
struct route_snapshot *snapshot;
u32 mtu = 0;
rcu_read_lock();
snapshot = rcu_dereference(active_route);
if (snapshot)
mtu = snapshot->mtu;
rcu_read_unlock();
return mtu;
}
This fragment assumes active_route is declared as an RCU-protected pointer and route_snapshot fields are immutable after publication. The returned scalar is copied while protection is active; returning snapshot itself would not be safe. If fields are intentionally modified in place, RCU alone is not their data-race protocol. Add a suitable lock, atomic operation, sequence counter, or another documented mechanism for those fields.
Normal RCU read-side critical sections cannot sleep. Do not call a potentially blocking operation, wait for work, or take a sleeping mutex while holding one. The exact preemption behavior depends on the configured RCU implementation, but the portable rule for normal RCU code is to keep the section short and nonblocking. SRCU is a different flavor for cases that need readers to sleep; it has a separately initialized domain and matching SRCU APIs, with different cost and lifetime rules.
The flavor is part of the reclamation contract. A reader protected by an SRCU domain must be paired with the corresponding SRCU grace-period or callback operation; synchronize_rcu() is not a universal wait for every RCU-like reader. Other specialized flavors serve task-exit or tracing requirements and have their own entry, exit, and grace-period APIs. Choose the narrowest documented flavor for the data structure, initialize and destroy its domain in the required lifecycle, and make the matching flavor visible in code review.
Publish a new version before retiring the old one
An updater first constructs a complete replacement privately. It then publishes the pointer with rcu_assign_pointer(). Readers that begin afterward can observe the new object; readers already using the old object can finish because the old allocation remains alive. A writer-side mutex below is not required by RCU itself, but it demonstrates how multiple update operations can be serialized.
struct route_snapshot {
struct rcu_head rcu;
u32 generation;
u32 mtu;
};
static DEFINE_MUTEX(route_update_lock);
static struct route_snapshot __rcu *active_route;
static void free_route_rcu(struct rcu_head *head)
{
struct route_snapshot *old;
old = container_of(head, struct route_snapshot, rcu);
kfree(old);
}
static int publish_route(u32 generation, u32 mtu)
{
struct route_snapshot *new, *old;
new = kmalloc(sizeof(*new), GFP_KERNEL);
if (!new)
return -ENOMEM;
new->generation = generation;
new->mtu = mtu;
mutex_lock(&route_update_lock);
old = rcu_dereference_protected(active_route,
lockdep_is_held(&route_update_lock));
rcu_assign_pointer(active_route, new);
mutex_unlock(&route_update_lock);
if (old)
call_rcu(&old->rcu, free_route_rcu);
return 0;
}
The code is an in-kernel API illustration, not a standalone userspace program. A real module needs the relevant kernel headers, error policy, and teardown path. The lockdep condition documents why this writer may inspect the old pointer without a read-side section: the update mutex serializes writers. RCU readers remain independent of that mutex. Never free the old snapshot immediately after rcu_assign_pointer().
The example deliberately treats each published snapshot as immutable. This copy-and-replace pattern makes the reader’s consistency model understandable: one traversal sees one valid object, even if a later version is published concurrently. If an update must change a linked list or tree in place, use the specific RCU list/tree helpers and their documented insertion, deletion, and traversal rules rather than inventing ordinary pointer stores.
Choose synchronous waiting or deferred callbacks
synchronize_rcu() blocks the updater until pre-existing normal-RCU read-side critical sections have completed. It does not wait for readers that begin after the grace-period wait starts. Use it when a sleepable updater needs to reclaim an object immediately afterward and can tolerate the latency. It cannot be called from an RCU read-side critical section or a context that cannot sleep.
call_rcu() registers a callback that runs after a grace period, letting the updater return without waiting. The callback must obey the restrictions of callback context; it should not block on a mutex or perform work that may sleep. For a simple RCU-protected allocation, kfree_rcu(object, rcu) can express deferred freeing without a custom callback. Do not schedule a callback for an object that remains reachable through an RCU pointer.
These APIs are not interchangeable with rcu_barrier(). A grace-period wait establishes that old readers are gone. A barrier waits for RCU callbacks that were queued before the barrier to finish. A module that queues callbacks whose functions live in that module must prevent unload until those callbacks have completed, typically by removing publications and then using the appropriate callback barrier in teardown. Follow the kernel subsystem’s lifetime requirements rather than substituting a generic wait.
Separate publication from writer consistency
RCU makes a read-mostly access pattern possible, but it does not resolve competing updates. Two writers that independently clone and publish a snapshot can overwrite one another’s changes unless they are serialized or use a compare-and-exchange protocol with retry. Allocation failure must leave the currently published object valid. Teardown must prevent new readers from discovering the object before it is reclaimed.
Pointer annotations and accessors are part of the contract. Declare protected pointers with __rcu, use rcu_dereference() in the matching read-side context, and use rcu_dereference_protected() only when the supplied condition proves another lock or invariant prevents concurrent updates. Use rcu_assign_pointer() for publication. A plain C assignment can omit the required ordering and make sparse or lockdep diagnostics less useful.
Grace periods are temporal, not reference counts. A reader does not increment a per-object count, and a grace-period wait does not identify which individual readers used an object. If the design needs arbitrary users to retain objects long after a lookup, a refcount or another ownership model may fit better. Combining RCU lookup with a safely acquired reference can be valid, but the handoff must be race-free and use the APIs documented for that object type.
Test lifetime boundaries, not only throughput
Review every pointer escape from the read-side section. Check that every allocation is fully initialized before publication, every old version is detached before retirement, and every callback can run while module state is still valid. A checklist should cover concurrent readers, concurrent writers, replacement during traversal, allocation failure, shutdown, and callback completion.
Use kernel diagnostics as complementary evidence: sparse catches address-space annotation mistakes, lockdep and RCU lockdep checks validate some context and pointer-use assumptions, and KASAN or KCSAN can expose lifetime or data-race defects under stress. The kernel’s RCU torture tests are valuable for validating RCU implementations; they are not a substitute for tests of a driver’s own object graph. A successful load or a fast read benchmark proves neither reclamation correctness nor update consistency.
Measure the workload after correctness is established. Compare reader latency and update cost with an appropriate lock-based design, include grace-period delay and deferred-callback backlog, and stress memory pressure. RCU can trade cheap readers for delayed reclamation and more complex ownership. That trade is worthwhile only when the access pattern justifies it and the lifetime proof remains reviewable.
Related:
- eBPF Explained: Safe, Programmable Observability in the Linux Kernel
- Tracing System Calls and Kernel Events with bpftrace
Sources: