Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux Driver Probe Deferral: Trace Supplier Dependencies and EPROBE_DEFER

Diagnose Linux driver probe deferral by tracing missing suppliers, resource lookup, deferred-probe reasons, dependency cycles, and retry behavior.

Linux device drivers do not always probe in the order that hardware designers expect. A consumer can depend on a clock, regulator, GPIO controller, reset controller, firmware provider, or another device whose driver has not bound yet. Returning -EPROBE_DEFER tells the driver core that a required dependency may appear later and the probe should be retried. It is a scheduling signal, not a generic error-recovery code and not a fix for a permanently missing resource.

The distinction between “not ready yet” and “will never be available” is central. A good diagnosis identifies the consumer, the precise resource it needs, the intended supplier, and whether the supplier can actually register. Repeatedly reloading modules or rebooting can hide ordering symptoms without repairing the dependency graph.

What the driver core does

During binding, the driver core invokes the matching driver’s probe callback. A return value of zero means binding succeeded. Other negative errors normally mean the driver did not bind and should have released resources it acquired. -EPROBE_DEFER is special: the core places the device on a deferred list and retries probing later as dependencies or driver registrations progress.

Drivers should defer as early as possible after identifying an unavailable mandatory supplier. Doing significant setup first means the same work must be unwound and repeated. The cleanup path must release acquired resources and leave no child devices registered. The driver documentation warns that returning -EPROBE_DEFER after creating child devices can produce repeated probe loops. Probe callbacks must therefore have explicit rollback behavior and should register children only after mandatory dependencies are ready.

Not every absent resource should cause deferral. An optional DMA engine may permit a PIO fallback; an optional reset line may have a documented no-reset mode. If the device can function in a supported degraded mode, the driver should represent that contract rather than wait indefinitely for a resource that may not exist. Conversely, silently substituting a dummy provider for a required clock or regulator can cause invalid hardware behavior.

Find the deferred consumer and reason

On a running system, collect device and driver state without changing bindings:

cat /sys/kernel/debug/devices_deferred 2>/dev/null
journalctl -k -b --no-pager | grep -i -E 'defer|probe|supplier|clock|regulator|gpio'
find /sys/bus -path '*/devices/*/driver' -type l -print 2>/dev/null

The devices_deferred debugfs file is available only when the relevant kernel support and debugfs mount are present. Its contents and reason detail depend on kernel version and driver use of the defer-reason helpers. An empty file does not prove that every device is healthy: a driver might have failed permanently with another error, or the debug interface may not be enabled.

For the consumer, record the bus path, modalias, bound driver if any, firmware node, and exact missing resource name. For the supplier, verify that a corresponding device exists, that its driver is built or installed, that it matched, and that its own probe succeeded. A supplier device can exist in sysfs while its driver is still unbound or failing.

Device links provide an explicit way to model certain supplier-consumer dependencies. Managed links can enforce both driver-presence and suspend/resume ordering; stateless links express ordering without requiring a supplier driver. They complement resource acquisition and firmware-description dependencies, but should not be added as a blanket workaround. The device-link documentation warns that a driver-presence dependency can defer a consumer indefinitely if the supplier is missing or blacklisted.

Trace resource lookup precisely

Follow the consumer’s probe path to the exact acquisition call. A failed regulator lookup, clock lookup, GPIO lookup, IRQ mapping, reset control, or NVMEM cell lookup can have different semantics depending on the API and whether the resource is optional. Read the driver’s handling of the returned error and the resource’s firmware binding. Do not infer the required name from a board schematic alone; the driver may request a specific connection ID.

For Device Tree systems, inspect the consumer node and supplier node, including phandles, names, status, and bus hierarchy. For ACPI, inspect the applicable resources and namespace path. Confirm that the provider node is enabled and that pinctrl or bus configuration lets the provider itself probe. A provider blocked by its own missing clock can create a dependency chain that reaches several devices.

Look for cycles. If driver A waits for B while B waits for A, each may remain deferred. A firmware dependency graph can also be incomplete or generated too late for the core to order correctly. The fix may be a corrected firmware description, a provider probe defect, or an inappropriate mandatory dependency, not a different module load order.

Distinguish delay from permanent absence

Boot-time ordering can legitimately cause transient deferral: a supplier driver registers after a consumer device was enumerated, then the core retries. A stable system should eventually bind the consumer when the supplier is ready. If the consumer remains in devices_deferred, test whether its supplier exists and inspect the stored reason or the driver logs.

A reason such as “supplier not ready” points toward supplier binding or an explicit dependency. A missing-resource error may indicate a misspelled property name, an omitted firmware node, or a resource that the platform never supplies. If logs report -EPROBE_DEFER indefinitely, identify which exact call returned it and whether the device core can know when to retry. A driver that defers for a condition that never generates another probe event can remain stuck until a separate reprobe occurs.

Do not use driver_override, unbind/rebind loops, or manual sysfs reprobe writes as the first test on a production host. Rebinding can interrupt devices, lose state, or trigger side effects. If reprobe is needed, use a lab system or an approved maintenance window and first confirm the driver’s remove/probe paths are safe.

Common root causes

Supplier driver is absent. The module may not be built for this kernel, may be blacklisted, may have failed signature or load checks, or may not match the supplier device. Confirm device presence, module availability, and supplier probe logs.

Firmware description is wrong. A phandle, supply name, clock name, interrupt, or GPIO property may point nowhere or use the wrong identifier. Compare the binding schema and the actual firmware description; do not patch the consumer to accept a different board’s wiring.

Provider itself cannot probe. The regulator or GPIO controller may depend on another supplier or have a pinmux, IRQ, or firmware error. Trace the chain upstream until reaching the first failing provider.

Optional resource treated as mandatory. A driver may return deferral when it should support a documented fallback. This is an implementation issue, not a board ordering problem.

Dependency graph cycle or late link. Managed device links can prevent probing until supplier drivers bind. A link created too late from the consumer probe requires the driver to verify supplier state and defer appropriately. Review the link flags and lifecycle against the kernel documentation.

Probe cleanup or repeated side effects. Deferral causes a probe retry. If the driver leaves regulators enabled, allocates duplicate child devices, or retains stale references after returning, retries can create secondary failures. A driver bug may be exposed by correct deferred probing.

Corrective workflow and acceptance

Use a focused sequence:

  1. Name the consumer device, driver, kernel release, and exact resource returning -EPROBE_DEFER.
  2. Identify the intended supplier and verify its device node, driver match, probe result, and dependencies.
  3. Validate the firmware binding and property spelling against the device’s authoritative binding and board description.
  4. Determine whether the resource is mandatory or optional and whether the driver handles it according to the hardware contract.
  5. Correct the actual missing provider, firmware description, or driver logic; avoid changing unrelated probe ordering.
  6. Reboot or reprobe only in an approved environment, then verify the consumer binds once and its functional interface works.

For acceptance, keep the before-and-after devices_deferred output, relevant boot log window, supplier and consumer sysfs paths, firmware-description revision, and actual hardware test. Confirm runtime suspend/resume and shutdown ordering if the devices have power dependencies. A device disappearing from the deferred list is necessary but not sufficient: validate the operation that originally required the supplier.

The safe interpretation is simple: deferral means “retry when a required dependency becomes available.” It does not mean “ignore this error,” “load modules until it works,” or “the consumer should wait forever.” Follow the dependency from the consumer to the first unavailable supplier, then fix that boundary and prove the full device path is functional.

Related:

Sources:

Comments