Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux CPUIdle: Target Residency, Wake Latency, and PM QoS

Understand Linux CPUIdle state selection by correlating target residency, exit latency, timer prediction, and power-management QoS constraints.

When a CPU has no runnable work, Linux can ask the processor to enter an idle state. Deeper states often save more energy but take longer to enter and exit. CPUIdle drivers describe the states available to logical CPUs, and governors choose one based on predicted idle duration and latency constraints. A wakeup that appears slow may be caused by a deep state, but it may also be timer placement, interrupt delivery, scheduling delay, firmware, or a device resume path.

CPUIdle is distinct from CPUFreq. CPUFreq selects an operating performance request while a CPU is executing; CPUIdle selects how the processor waits when a CPU is idle. Both influence power and latency, but changing a frequency governor does not directly define idle-state selection.

Driver-supplied idle states

The CPUIdle driver provides an abstract list of states for each CPU. Each state is characterized by properties including target residency and worst-case exit latency. Target residency estimates how long hardware must remain in a state, including entry cost, for it to save more energy than a shallower state. Exit latency estimates how long it takes before the CPU can execute after a wakeup, including the case where a wakeup occurs while the processor is entering the state.

These values are driver/platform descriptions, not universal constants. Firmware settings, microcode, kernel version, and processor family can change which states are available or how they are named. Check per-CPU state directories and counters rather than assuming state3 means the same thing across models.

for state in /sys/devices/system/cpu/cpu0/cpuidle/state*; do
    [ -d "$state" ] || continue
    printf '\n%s\n' "$state"
    for attr in name desc latency residency usage time disable; do
        [ -r "$state/$attr" ] && printf '  %s=' "$attr" && cat "$state/$attr"
    done
done

The exposed attributes depend on the kernel and driver. Read the relevant sysfs ABI and compare counters over a time interval. A cumulative usage count or time counter cannot explain a single latency event without a baseline and wakeup timeline.

Governor decision and prediction

The governor receives an expected sleep length and considers idle-state target residency and exit latency. The predicted idle duration is uncertain: a timer or interrupt may arrive earlier than expected. Governors may refine their decision based on recent idle history and whether the scheduler tick can stop. The deepest state is not always the best choice, especially for short idle windows or workloads with strict wake latency.

Power-management QoS can impose a maximum acceptable wakeup latency. The governor should avoid states whose exit latency exceeds the applicable constraint. A device driver or userspace process can therefore influence idle-state decisions indirectly by requesting a latency bound. If a platform never selects a deep state, inspect active QoS requests before assuming the CPUIdle governor is malfunctioning.

The scheduler tick also affects predictions. If the next timer is near, the CPU may not remain idle long enough to amortize a deep state’s entry and exit costs. Tickless operation can allow longer idle intervals when no timer requires regular scheduler activity, but it does not eliminate device interrupts or guarantee that the predicted duration will occur.

Measure actual wake latency

Distinguish hardware exit latency from application response time. A userspace event handler must wait for interrupt delivery, kernel processing, task scheduling, and application execution after the CPU begins leaving its idle state. A high observed request latency does not prove that the idle state itself exceeded its documented latency.

Use kernel tracing to compare the wakeup source, interrupt timestamp, CPUIdle state entry/exit, task wakeup, and task execution. Test with the same interrupt source and workload at idle and under load. A scheduler delay or interrupt storm can dominate the time measurement even when the hardware exits promptly.

Measure power and latency together. Disabling all deep states may reduce wakeup latency but increase idle energy and heat. On mobile or embedded systems, an apparently small per-event increase in active power can dominate battery life across many idle intervals. Determine an acceptable latency budget before tuning.

Common diagnostic traps

Do not map a human-readable state name from one processor to another without platform evidence. Do not infer a state from residency counters alone if the counters are unavailable, disabled, or reset at a different time. Some states are coordinated across cores or packages, and a single CPU’s view may not show the full hardware decision.

A PM QoS constraint can come from a device driver or active application, so a broad search for latency in sysfs may not reveal the requester. Use the supported PM QoS debug or tracing facilities for the target kernel. Avoid writing a global latency request as a permanent fix without identifying the owner and lifetime; a stale constraint can prevent energy-efficient states indefinitely.

Latency constraints should represent a real service requirement. A driver that requests an unnecessarily small maximum wakeup latency may keep CPUs in shallow states across the entire system, while an absent constraint can allow a state that violates an application’s response budget. Measure the event-to-action deadline and include device, interrupt, scheduler, and application time before setting the limit. Release the request when the consumer is inactive so a temporary foreground workload does not govern power behavior forever.

Idle-state accounting is also per CPU in many interfaces. A CPU that receives periodic interrupts can show little deep-state residency even if the package is mostly idle. Another CPU may enter a deep state while a sibling remains active, depending on processor topology and hardware coordination. Capture per-CPU activity together with package-level power data where available; do not average per-CPU residency and call it a package state.

Firmware can also disable or limit states. A kernel command-line parameter that hides a state can be useful for a controlled comparison, but it should not be left in production without measured justification. A state that causes a platform-specific hardware problem should be reported with processor model, BIOS/UEFI revision, kernel version, and repeatable reproduction steps.

Controlled validation

Use a test matrix with idle, short periodic wakeups, long idle intervals, and a representative latency-sensitive workload. Collect state usage and time deltas, power consumption, wakeup latency distribution, temperature, and workload performance. Repeat on AC and battery or on relevant power profiles, since firmware policy can vary.

Run before/after tests under the same kernel and workload, changing only one state or QoS source at a time. Confirm that the desired state remains available after reboot and that test tooling has not left a latency request or kernel parameter behind. Validate both the target latency and energy cost over a representative operational period.

Include timer-heavy and interrupt-heavy tests as well as long idle periods. A periodic heartbeat can keep the scheduler tick or a device interrupt active and shorten predicted sleep intervals. If the goal is lower idle power, locate the wakeup source before changing state limits. If the goal is lower latency, identify the path that consumes the budget rather than tuning the maximum exit-latency parameter to compensate for unrelated userspace delay.

State disable controls, when exposed, are useful for a controlled comparison but should not become an undocumented production setting. Disabling a state can alter package coordination and shift residency into another state, so the resulting power effect may not be monotonic. Capture the before/after residency distribution and battery or wall-power measurements instead of assuming that excluding one state saves energy.

On heterogeneous systems, different CPU clusters may expose different state lists and latency properties. A workload migrating between clusters can therefore see different wakeup behavior even with the same kernel policy. Record CPU affinity, topology, and scheduler placement during tests. If latency changes when a task is pinned, inspect both the CPUIdle state data and the interrupt routing path before attributing the difference to one mechanism.

CPUIdle is a prediction-and-constraint system. A driver describes the idle states, a governor chooses among them, and PM QoS limits which latencies are acceptable. Diagnose the complete wake path, not only the deepest state name, and balance measured energy savings against the workload’s actual response-time requirement.

Related:

Sources:

Comments