Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux Device Runtime PM: Usage Counts, Autosuspend, and Callback Order

Implement Linux device runtime power management with balanced usage references, autosuspend policy, parent-child dependencies, and observable recovery.

Runtime power management allows a device to enter a low-power state while the system remains operational. The PM core tracks activity, coordinates callbacks, and can apply an autosuspend delay after the device becomes idle. The driver must still define what “idle” means, balance every usage reference, restore hardware state before I/O, and respect parent-child dependencies. A device that never autosuspends can waste power; a device suspended while a request is in flight can fail intermittently.

This article covers the Linux device runtime PM framework. It is separate from CPUIdle and system suspend, although runtime PM callbacks interact with system-wide sleep and resume transitions. Follow the current kernel documentation for the target release because helper availability and subsystem behavior can evolve.

Usage references and runtime status

The runtime PM usage counter prevents a device from suspending while a consumer holds an active reference. A driver or subsystem should acquire a runtime PM reference before accessing hardware that may be suspended, then release it when the device is no longer needed. The counter is not a general object reference count; it represents active hardware usage and must be paired on every success and error path.

For synchronous access, pm_runtime_resume_and_get() resumes the device if needed and increments the usage count on success. When done, an appropriate pm_runtime_put_*() helper releases the reference and may trigger idle or autosuspend. If the resume operation fails, do not proceed as though the hardware is active.

static int read_status(struct device *dev)
{
    int ret = pm_runtime_resume_and_get(dev);

    if (ret < 0)
        return ret;

    ret = perform_status_read(dev);
    pm_runtime_mark_last_busy(dev);
    pm_runtime_put_autosuspend(dev);
    return ret;
}

The helper names and usage pattern must be checked against the target kernel headers and subsystem. If perform_status_read() can fail, the reference still must be released. In production code, structure cleanup with a single exit path or a scoped helper so no error branch leaks the usage count.

The distinction between pm_runtime_get_sync() and pm_runtime_resume_and_get() matters for error handling: the former historically increments the usage counter even when the resume operation fails, so callers must pair the reference appropriately; the latter combines resume and increment with a clearer success contract. Follow the current API documentation and inspect the subsystem’s established pattern. A leaked reference tends to manifest as a device that never suspends, while an extra put can lead to an invalid counter or premature power-down.

Make every path explicit: normal completion, short-circuit validation, I/O error, timeout, interrupt cancellation, and device removal. If a request is asynchronous, ownership of the runtime PM reference must last through the callback that actually touches hardware. Releasing it when the request is queued instead of when it completes creates a race where autosuspend begins while DMA or register access is still active.

Callbacks and hardware state

Subsystem callbacks typically implement runtime suspend and runtime resume for a device. Suspend should place hardware in the supported low-power state and report errors if that transition is unsafe or incomplete. Resume should restore required registers, clocks, power rails, and internal driver state before new I/O proceeds. The callback context and locking rules depend on the PM core and subsystem; do not sleep or take locks in a way that conflicts with the callback’s documented context.

The driver’s idea of device state can diverge from physical state. Firmware may reset a device during system sleep, a parent may power-cycle it, or a runtime resume can race system resume. The PM core may bring devices to full power during system resume and expects subsystem state to be reconciled. Runtime callbacks should be idempotent where the hardware contract requires it and should distinguish already-active state from a genuine failure.

Autosuspend policy

Autosuspend inserts a delay between the last activity and a runtime suspend request. The driver or subsystem enables autosuspend and marks the last busy time after operations. A non-negative delay can be adjusted through device power attributes by userspace after registration, so the value may be platform policy rather than a compile-time constant.

Use autosuspend when short idle gaps are common and entering low power has a meaningful cost. Too short a delay can cause repeated suspend/resume churn, device errors, and latency. Too long a delay can waste energy. Measure idle patterns, transition cost, wake latency, and workload behavior before choosing a default. Do not hard-code one delay as universally optimal for every device.

The autosuspend helpers call the documented last-busy mechanism and schedule a suspend when appropriate. Mixing autosuspend and non-autosuspend helpers without a clear policy can make behavior difficult to predict. Select one usage pattern and document the delay owner.

An autosuspend delay of zero does not eliminate every race or guarantee an immediate hardware transition. The PM core may defer work, the parent may still be active, a wakeup may arrive, or a callback may reject the transition. The usage counter and runtime status describe the framework’s state; the driver must still keep that state synchronized with the device. Monitor callback return values and transition statistics instead of assuming a put equals “power off now.”

Parent-child relationships

A child device may depend on its parent being powered. Runtime PM tracks device hierarchy and can coordinate parent state, but drivers still need to describe their dependencies correctly. A parent that suspends while a child is active can make the child inaccessible. A child that holds an unnecessary reference can keep an entire hierarchy powered.

Test probe, remove, hot-unplug, runtime idle, and parent resume paths. If a bridge or controller is required for child resume, ensure the PM ordering and reference behavior reflects that dependency. Do not add a global pm_runtime_get_sync() call as a workaround without pairing it and explaining which device lifetime requires the reference.

System sleep interaction

Runtime PM work is coordinated with system-wide suspend and hibernation. The kernel documentation recommends using pm_wq for work items related to runtime PM so they can be synchronized with global power transitions. If a runtime-suspended device is restored to full power during system resume, its PM-core runtime status must match the actual hardware state. A stale runtime status can cause the next request to skip a required resume callback.

Test transitions while the device is active, idle, in the middle of autosuspend, and after an interrupted request. Include a child-device hierarchy and at least one failed resume. Verify that system suspend can quiesce pending runtime PM work and that device state is correct after resume.

Diagnose a device that will not suspend

Inspect the device’s runtime status, usage count, autosuspend delay, last-busy time, and error fields where exposed under its power directory. The exact sysfs attributes depend on configuration and device type. A positive usage count often indicates a consumer reference remains held; repeated busy transitions can indicate a workload or autosuspend delay; a recorded runtime error points to a callback or hardware transition failure.

Do not forcibly write a device to suspended to hide a driver bug. First identify the consumer, trace reference increments/decrements, and examine suspend callback logs. Compare system-level energy and device functionality before and after a supported fix.

Validation and review

Use tracepoints and PM debug messages to correlate usage references, callback entry/exit, and hardware transitions. Test under I/O load, idle, unplug, error recovery, and system sleep. Inject failures where the subsystem supports it. Confirm every resume-and-get has a matching release, including error paths, and that the active state always matches hardware.

For a driver review, draw the active-to-suspended state machine and annotate which lock protects each hardware transition, which callers hold usage references, and what event marks the last activity. Include the parent device’s state and any remote wake mechanism. A review that only searches for pm_runtime_enable() and pm_runtime_put() can miss an asynchronous work item that outlives the reference or a resume callback that does not restore cached register state.

Runtime PM is a state machine with reference accounting, transition callbacks, policy delay, and hierarchy. Use balanced references, choose autosuspend based on measurement, synchronize with system transitions, and test the actual hardware lifecycle rather than assuming a zero usage count proves the device is safely off.

Related:

Sources:

Comments