Linux Kernel Workqueues: Concurrency, Ordering, and Forward Progress
Choose Linux workqueue attributes safely by reasoning about sleepability, concurrency limits, reclaim dependencies, flushing, cancellation, and forward progress.
Kernel code often needs to defer work that cannot run in its current context. A network callback, interrupt handler, or state transition may need to schedule operations that can sleep, perform bounded cleanup, or interact with a subsystem later. Linux workqueues provide a mechanism for queueing such work for execution by worker contexts. The difficult part is not calling queue_work(): it is selecting the right execution domain, defining ownership and cancellation, and ensuring dependencies cannot prevent forward progress.
This article discusses the contemporary concurrency-managed workqueue design described by the Linux kernel documentation. Workqueue internals and flags evolve; verify the target kernel’s Documentation/core-api/workqueue.rst and header API before using an example in a driver. Workqueues are not a universal replacement for softirqs, tasklets, threaded IRQs, or a dedicated kernel thread. Choose the execution mechanism from the required context and latency.
Work item versus workqueue
A work item represents a deferred callback and its state. A workqueue represents a domain for attributes, ordering, flushing, and forward-progress guarantees; modern concurrency-managed workqueues share worker pools rather than creating a private thread for every queue. A bound workqueue normally has CPU locality, whereas an unbound workqueue is not tied to the CPU that queued the item. Flags can change locality, priority, reclaim, and ordering behavior.
The distinction matters when estimating concurrency. A dedicated workqueue is not necessarily a dedicated worker thread. If a work item is queued twice while it is already pending, the queueing call does not create an independent parallel copy of that same work structure. If a producer needs to represent multiple outstanding events, it must keep a separate counter, list, or work item per logical request. Do not assume a work item is a lossless event queue.
struct sample_device {
struct work_struct recovery_work;
struct workqueue_struct *recovery_wq;
atomic_t recovery_requested;
};
static void recovery_workfn(struct work_struct *work)
{
struct sample_device *device =
container_of(work, struct sample_device, recovery_work);
/* This context may sleep; it still must bound I/O and retry behavior. */
recover_device_state(device);
}
static int sample_init(struct sample_device *device)
{
INIT_WORK(&device->recovery_work, recovery_workfn);
device->recovery_wq = alloc_workqueue("sample-recovery", 0, 0);
return device->recovery_wq ? 0 : -ENOMEM;
}
This sketch omits device-specific error handling and teardown. The callback’s ability to sleep does not make sleeping forever acceptable. It should use bounded waits, state checks, and clear completion rules so a flush or device removal can finish.
Context and sleepability
Ordinary threaded workqueue callbacks run in worker process context and may sleep where the called APIs allow it. BH work items are different: they run in the queueing CPU’s softirq context, cannot sleep, are per-CPU, and have tighter API restrictions. Selecting a workqueue based only on the desired name or priority can therefore violate context rules. Trace the full callback chain and confirm every function is valid in the selected context.
Deferring from an interrupt does not transfer ownership automatically. The producer must make any shared state safe before queueing the work. If a device can be removed between the interrupt and callback, synchronize teardown so the callback cannot dereference freed state. A queued work item can outlive the function that queued it; its containing object must remain allocated until the work has completed or has been synchronously canceled.
Avoid using a workqueue as a substitute for a direct, short state update if a lock or atomic operation is sufficient. Excessive deferred work increases latency and obscures the causal sequence. Conversely, do not run potentially blocking operations from atomic context and hope that a timing difference makes them safe.
Concurrency limits, ordering, and dependencies
max_active limits how many work items may be active for a workqueue. The effective limit has topology-specific behavior: for bound queues it applies per CPU; for unbound queues the kernel distributes concurrency by NUMA node. Interdependent work items can deadlock if a dependency requires another item to run but all available active slots are occupied by items waiting on that dependency. The kernel documentation recommends the default (0) unless the caller has a specific throttling or ordering requirement.
For strict single-flight ordering, use an ordered workqueue or an appropriate single-active configuration supported by the target kernel. Do not infer global ordering from a queue name or from observed behavior on one CPU. If work on two queues communicates, their independent worker pools do not create a dependency ordering. Use an explicit completion, lock, or state transition that establishes the intended relation.
Reentrant work must also be considered. A callback that synchronously flushes or waits for another item on the same constrained queue can block the worker needed to satisfy the wait. Draw a dependency graph for producers, work items, locks, completions, and flush points. If there is a cycle, changing a timeout or increasing concurrency may only make the failure harder to reproduce.
Queue attributes are part of that execution contract. High-priority workqueues still participate in worker-pool behavior and should not be used as a substitute for a latency budget. CPU-intensive callbacks have different interactions with concurrency management than callbacks that frequently sleep. Unbound queues are useful when CPU locality is not required, but they change placement and concurrency assumptions; they are not automatically faster. Inspect the current kernel’s workqueue_attrs interface and documentation instead of carrying forward a deprecated queue-creation pattern from an old driver.
For a stream of requests, the queue’s concurrency setting is not backpressure at the producer boundary. If requests can arrive faster than callbacks complete, track queue depth and define whether to coalesce, drop, reject, or persist work. A work item that repeatedly requeues itself without a limit can monopolize service or keep teardown busy indefinitely. Make retry policies bounded and expose a counter or trace event so operators can distinguish a slow device from a producer that never stops scheduling work.
Forward progress during memory reclaim
Some work must run to free memory or release resources while the system is reclaiming memory. Such a workqueue may require WQ_MEM_RECLAIM, which gives the queue a reserved execution context (rescuer) to support forward progress under constrained worker creation. Omitting the flag for reclaim-critical work can form a dependency cycle: reclaim waits for the work item, but the worker needed to run the item cannot be created without memory.
The flag is not a general priority boost and does not fix arbitrary circular dependencies. If several reclaim-dependent operations rely on one another, the kernel guidance may require separate reclaim-capable queues. Trace which item actually breaks the cycle. Do not mark every queue reclaim-capable without understanding the memory dependency graph.
Flush, cancel, and destruction semantics
flush_work() waits for a specific work item to finish; flush_workqueue() waits for work that was queued on the queue when the flush began. New incoming work does not livelock that flush, but may fall outside its completion boundary. A flush is a synchronization point, not a request to stop producing work or proof that the queue is empty when it returns. For teardown, stop and synchronize producers first, then flush or cancel the relevant work; ensure no racing queue calls remain.
cancel_work_sync() prevents pending work from running and waits for an in-progress execution to finish. It must not be called while holding a lock that the callback itself needs, or teardown can deadlock. Delayed work has its own cancellation requirements. Destroy the queue only after callbacks no longer need the object and all relevant work items are drained. Use device-managed allocation only when its teardown order matches the driver’s lifetime semantics.
Diagnose and test with lifecycle evidence
Workqueue stalls may appear as a hung remove path, delayed recovery, memory reclaim pressure, or a system-wide wait chain. Add tracepoints or subsystem instrumentation around queueing, start, finish, and cancellation, keyed by a device or request identifier. Do not log high-volume payloads or infer causal order solely from wall-clock timestamps. Kernel trace events can show worker-pool activity and help identify queue saturation or dependency chains.
Test repeated queueing, producer shutdown, cancellation while running, device removal under load, allocation failure, memory pressure, CPU hotplug if relevant, and a work item that queues follow-up work. Verify that each event is either coalesced intentionally or represented independently. The workqueue design is sound when every callback’s context, lifetime, maximum concurrency, ordering, and forward-progress needs are explicit.
Use lockdep and kernel debugging facilities in a representative development configuration where possible. A concurrency test should create competing producer and teardown paths, then repeat the sequence enough times to expose a rare queueing race. Measure both queue delay and callback duration; a long callback and a starved callback require different fixes. Preserve kernel version and configuration with the trace because worker implementation details and available attributes vary across releases.
Related:
- Linux NAPI Receive Processing: Poll Budgets, Queue Scaling, and Latency
- Linux RCU in Practice: Read-Side Sections, Grace Periods, and Reclamation
Sources: