Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux CPUFreq: Policies, Governors, and What Frequency Readings Mean

Diagnose Linux CPU performance scaling by separating CPUFreq policies, governors, driver requests, hardware limits, and measured workload behavior.

Linux CPU frequency scaling is a coordination problem between the scheduler, a scaling governor, a platform driver, firmware, and the processor. A sysfs value labelled “current frequency” may be the last performance request rather than a live measurement of the clock actually running on the hardware. A policy may also cover several CPUs that share one control interface. Reliable diagnosis starts by identifying the active driver and policy topology, then comparing requests with independent performance and thermal evidence.

The kernel CPUFreq documentation describes a core, scaling governors, and hardware-specific drivers. Some drivers can use hardware feedback and implement their own scaling algorithm rather than attaching a generic governor. Therefore, a tuning guide for one driver cannot be applied blindly to another. Read the active policy and scaling driver on the target system before changing a value.

Policy objects group CPUs that share control

CPUFreq represents CPUs that share a hardware performance-scaling interface in a cpufreq_policy. A policy can cover one CPU or several CPUs. Per-CPU sysfs links can point to the same policyX directory. This means “set the frequency for CPU 3” may really mean “change the performance bounds for a policy shared by CPUs 2 through 5.” Inspect related_cpus and affected_cpus before interpreting any policy setting.

The policy interface lives under /sys/devices/system/cpu/cpufreq/policyX/ on systems where the driver exposes the generic sysfs attributes. Not all systems provide every file. Firmware can impose a limit, driver-specific attributes may exist, and available governors depend on the selected driver and kernel configuration.

for policy in /sys/devices/system/cpu/cpufreq/policy*; do
    [ -d "$policy" ] || continue
    printf '\n%s\n' "$policy"
    for attr in related_cpus affected_cpus scaling_driver scaling_governor \
                scaling_min_freq scaling_max_freq scaling_cur_freq bios_limit; do
        [ -r "$policy/$attr" ] && printf '  %s=' "$attr" && cat "$policy/$attr"
    done
done

This is read-only inspection. Frequency values are commonly expressed in kHz by the generic policy interface, but verify the ABI on the target kernel. An absent file means the driver does not expose that attribute; it is not equivalent to zero or unlimited.

Governor request is not physical measurement

A governor estimates the performance level needed by a policy. The driver then asks the platform to change a performance state. Firmware, hardware coordination, thermal protection, current limits, and the processor’s own behavior can affect the actual result. scaling_cur_freq often describes the last requested P-state, not an instantaneous clock sample. Some architectures provide more precise readings under additional conditions, but the attribute still may not exactly equal the physical frequency at the moment it is read.

This distinction matters when a benchmark looks slower than a reported frequency suggests. The scaling request can be high while thermal or power limits reduce effective work. Conversely, a fast turbo interval can end before a low-rate sampling tool records it. Compare sustained throughput and latency, not only one frequency counter.

Generic governors implement policy-level algorithms. performance requests the highest frequency within the maximum policy limit; it does not overrule the limit. powersave requests the lowest allowed frequency for that policy, but hardware-specific drivers can interpret governor semantics differently. schedutil uses scheduler utilization information. A driver such as intel_pstate can provide driver-specific algorithms and may bypass the generic governor layer. Read the driver documentation before assuming two policy names behave identically across platforms.

Limits and other controllers

scaling_min_freq and scaling_max_freq bound requests for a policy. Setting them does not guarantee exact operating points. Firmware may expose a bios_limit that caps CPUFreq independently. Thermal limits may be managed through the generic thermal subsystem and are not necessarily represented by bios_limit. A platform can also impose power or current limits that alter the result.

Changing a policy on a production host can affect every CPU in its policy group. Do not set minimum or maximum values from a generic internet example. First record the original bounds, governor/algorithm, driver, CPU mask, kernel command line, thermal behavior, and power profile. If an experiment is necessary, change one variable on a controlled system, use the supported management layer, and define a rollback.

The nominal frequency table may also be a poor proxy for delivered capacity. Modern processors can change clock rapidly, use boost ranges, or expose an abstract performance scale rather than a fixed frequency. An operating point requested by the kernel is not a guarantee of a particular instruction rate because memory stalls, cache misses, SMT contention, and workload vector behavior all affect useful work. Compare work completed per unit time and the relevant hardware counters where the platform exposes them.

If several CPUs share a policy, a per-core utilization graph can be misleading. A busy CPU can raise the shared policy request and affect an otherwise idle sibling. Conversely, a shared hardware control interface can constrain what an isolated CPU can request. Check CPU topology, affinity, and policy membership together. Do not infer that one thread owns the observed frequency simply because it was pinned to a particular logical CPU.

Diagnose a performance complaint

Begin with the workload’s actual symptom: lower throughput, longer latency, inconsistent performance, or high power draw. Record CPU utilization, run queue, frequency policy, thermals, power profile, and benchmark duration. A short benchmark can measure burst behavior while a sustained workload measures thermal equilibrium. If the workload is I/O-bound, CPUFreq may not be the limiting factor at all.

Compare the governor request with an architecture-appropriate hardware measurement and with performance counters where available. Keep the measurement interval and sampling overhead in mind. Correlate it with thermal-zone and cooling-device state, firmware events, and power-limit counters. When using virtualization, the guest may see a virtual performance interface or counters whose semantics are controlled by the hypervisor.

Check whether userspace power-management software is rewriting policy limits or governor selections. A one-time sysfs change may be overwritten by a daemon or policy service. Identify the source of the active setting before editing it; otherwise the system can appear to revert spontaneously after login, AC/battery changes, or a device event.

Validation and rollback

Use the same workload, CPU affinity, temperature baseline, and power source for before/after comparisons. Warm up the system, run repeated samples, report distributions, and monitor thermal behavior through the full test. Validate idle power as well as peak performance; a maximum-frequency request can create more heat and reduce sustained capacity.

For latency-sensitive services, measure tail latency and throughput at the same time. A policy that improves the first burst may raise package temperature and lower sustained performance later. For batch jobs, total completion time and energy-to-completion can be more informative than maximum observed clock. Keep background workloads controlled so scheduler utilization does not cause a different governor response across test runs.

Frequency reports from monitoring tools should identify their source. A tool may read the generic CPUFreq policy, an architecture-specific performance counter, or a scheduler capacity estimate. These values answer different questions and may sample at different rates. Before comparing two dashboards, inspect the underlying file or kernel API and confirm that both represent the same policy, unit, and measurement semantics.

If the CPUFreq driver is absent, the kernel may still boot and process work, but policy attributes and scaling governors will not be available as expected. Check the driver binding and kernel configuration rather than repeatedly writing nonexistent sysfs files. Driver initialization can fail because firmware tables, platform mode, or hardware support differ from what a generic guide assumes.

Record the original state so it can be restored. After changing a policy, check every CPU linked to it and ensure no other policy was affected unexpectedly. Reboot or restart the power-management service if required by the platform, then confirm the expected state persists only if persistence is intended. Do not edit firmware or kernel boot parameters as a first-line response to a workload regression.

CPUFreq is best understood as a policy request path, not a direct clock dial. Inspect the CPUs sharing each policy, identify the active driver and scaling algorithm, distinguish request from measurement, and account for thermal, firmware, and workload limits before making changes.

Related:

Sources:

Comments