Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux lockdep: Validating Lock Ordering and Interrupt Contexts

Use Linux lockdep to detect exercised lock-order and interrupt-context hazards, interpret reports, and build actionable kernel locking tests.

Locking bugs are difficult because a deadlock may require a particular sequence of otherwise valid operations. Linux lockdep is a runtime validation framework that tracks lock classes and observed acquisition dependencies, then reports when exercised paths create a potential dependency cycle or violate interrupt-context rules. It can expose a dangerous ordering before the exact deadlock occurs. It is not a proof that the kernel is free of locking bugs: it can only reason about the lock behavior and paths that the configured kernel and workload actually exercise.

The most useful way to treat lockdep is as an executable consistency check for locking rules. A report is evidence that the kernel observed a suspicious relationship. The next task is to understand the locking contract, not to silence the warning. A real fix may change acquisition order, shorten a critical section, use a different primitive, or accurately annotate intentional nesting. Suppressing diagnostics without reviewing all call paths can convert a visible failure into an intermittent production deadlock.

What lockdep models

Lockdep groups lock instances into classes, generally according to lock type and initialization context. When a task acquires one class while holding another, the framework records an ordering edge. If another observed path establishes the reverse order, the resulting dependency graph can contain a cycle. The cycle indicates a possible deadlock if two execution paths acquire the locks concurrently in opposing order; it does not require the deadlock itself to have happened during the test.

The class model explains why a lockdep splat often contains a chain rather than a single “bad lock.” Read the full chain of acquisition sites and the held-lock state. Two source lines that look unrelated may operate on different instances of the same class; conversely, a lock subclass annotation can distinguish legitimate nesting when the locking design truly has a hierarchy. The annotation should describe the design, not merely make a warning disappear.

Lockdep also checks relationships with interrupt context. A lock acquired in process context and then acquired in hard-interrupt context can be unsafe if the interrupt can preempt the same CPU while the process holds the lock. The process can spin waiting for an interrupt handler that cannot run to completion until the process releases the lock. IRQ-safe locking rules and lockdep’s usage-state tracking help identify such inversions. The exact configuration and reporting depend on kernel options and architecture.

Build or obtain a kernel that can validate locking

Lockdep support is controlled by kernel configuration. Development and test kernels commonly enable lock dependency validation and related proving options, but distribution kernels differ. Check the running configuration when available:

grep -E 'CONFIG_(LOCKDEP|PROVE_LOCKING|DEBUG_LOCK_ALLOC|DEBUG_LOCKS)' \
  /boot/config-"$(uname -r)" 2>/dev/null

Some systems expose the active configuration through /proc/config.gz; others do not ship it. A missing line in a file is not proof that a facility is disabled unless the configuration source is known to be complete. Confirm the build configuration and kernel documentation for the exact release. Enabling debugging options can add runtime overhead and memory use, so use an appropriate test kernel for broad stress workloads.

Do not change kernel configuration on a production host merely to clear a diagnostic. A lockdep report can already be present in the kernel log, and runtime disable controls are not a substitute for capturing the original splat. Preserve the complete message, kernel release, configuration, workload, and reproducible operation before restarting or changing debug settings.

Reproduce with coverage, not just load

Lockdep learns from executed acquisitions. A long stress test that repeatedly follows one path may miss a rare path that takes the same locks in reverse order. Exercise both code paths, error handling, callbacks, teardown, suspend/resume where relevant, and interrupt-driven activity. Tests should cover resource creation and destruction because cleanup often has a different lock order than steady-state operation.

Use a test matrix that maps public operations to relevant locking paths. For a driver, this can include probe, open, I/O completion, reset, timeout, hot removal, and error recovery. For a core subsystem, it may include normal operations plus concurrent teardown. Run kernel selftests or subsystem tests where available, then supplement with workloads derived from the report. Keep the test bounded and record exact steps so the lock graph can be compared across revisions.

A clean run means only that no checked violation was observed under that configuration and workload. It does not show that every possible lock order is safe. Code review should still reason about the intended hierarchy and concurrent execution. Static analysis, targeted stress, and lockdep complement one another; none replaces the others.

Read a report systematically

Start with the headline and the first acquisition trace. Identify the lock classes, whether they are nested, the owning task, and any interrupt-state information. Then follow the dependency chain in order. Reports often include the previously observed acquisition site and the current path that closes the cycle. Compare the two paths against the source code and document which locks can be held concurrently.

Ask these questions before editing:

  1. Are the locks truly instances of the same class, or does initialization need a justified subclass?
  2. Can the two reported paths run concurrently, including through callback or workqueue execution?
  3. Does one path hold a lock while waiting for a worker, completion, or interrupt that needs that lock?
  4. Is the lock acquired in process, softirq, or hard-IRQ context, and are the relevant save/restore rules correct?
  5. Does an error or teardown path reverse the normal acquisition order?

Lockdep reports can be triggered by a path that is technically unreachable under a subsystem invariant, but that conclusion requires evidence. Explain the invariant, show how it is enforced, and test the boundary. Do not label a warning “false positive” solely because the deadlock has not reproduced.

Repair the locking design

The most robust correction is usually a consistent global ordering. If all paths acquire lock A before lock B, a reverse acquisition should be changed or redesigned. Sometimes the apparent critical section is too broad: copying data under a lock and performing slower work afterward can reduce nesting, provided the object lifetime and state remain protected. Replacing a lock with a different primitive is not automatically safer; the new primitive has its own ownership and context rules.

Nested lock subclasses are appropriate when the same lock class is intentionally acquired in a strict hierarchy, such as parent then child. The subclass must encode that hierarchy at each acquisition site and must match all possible nesting. It does not make arbitrary same-class recursion safe. If two objects can be locked in either order, sort them by a stable ordering before acquisition or use another design that avoids opposite nesting.

For interrupt-context problems, choose the right IRQ-safe variant only when the lock is actually shared with the relevant interrupt context. spin_lock_irqsave() saves and disables local interrupts around acquisition; it is not a universal cure for all context issues and has performance and latency implications. Analyze softirq and hard-IRQ paths separately, and use the primitive appropriate for the kernel context and configuration. Consult subsystem conventions before altering established locking patterns.

Verify the fix and preserve evidence

After a correction, reproduce the original test with lockdep enabled, add regression coverage for both acquisition orders, and run relevant subsystem tests. Check that the exact warning is gone, but also inspect neighboring reports: a changed order may reveal a different cycle or expose a missing annotation. Measure performance and latency if the fix broadens IRQ-disabled regions or adds serialization.

Keep the complete before-and-after splats with kernel versions, configuration, commit identifier, and workload. Record whether a report was fixed by ordering, scope reduction, lifetime restructuring, or a justified subclass. That context prevents a later refactor from restoring the hazard and helps distinguish a genuinely new cycle from a known, evidence-backed constraint.

Operational acceptance should be specific: the intended paths run under a lockdep-capable kernel; the original cycle no longer appears; coverage includes both normal and teardown/error paths; no new lockdep warnings are introduced; and code review documents the lock hierarchy. A single boot with no splat is not an adequate acceptance test for a concurrency fix.

Lockdep is valuable because it turns lock ordering assumptions into runtime-checkable evidence. Its output is neither noise to suppress nor a mathematical proof. Build a lock hierarchy, exercise the paths that can violate it, analyze the entire report, and verify the repair with repeatable tests.

Related:

Sources:

Comments