Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux Thermal Sysfs: Trace Zones, Trip Points, and Cooling Devices

Diagnose Linux thermal behavior by correlating thermal zones, trip points, cooling-device states, hwmon values, and workload symptoms without blind tuning.

When a Linux system slows under load, temperature alone does not prove thermal throttling. The kernel thermal framework represents sensors and cooling actuators as thermal zones, trip points, and cooling devices. A zone can associate a sensor with one or more actions such as fan speed or processor cooling. Hardware drivers may also expose sensor readings through hwmon. The operational task is to correlate these interfaces with workload and performance evidence, not to change thresholds until the symptom disappears.

The exported interface is platform-dependent. Not every machine registers every sensor, governor, trip point, or cooling device. Names such as thermal_zone0 are enumeration IDs, not stable descriptions of a CPU core. Read each zone’s type, available policy, trips, and linked cooling devices before making an interpretation.

Thermal zone model

A thermal zone describes a temperature source and policy relationship. The sysfs tree can expose a current temp, mode, policy, supported policies, and trip-point temperatures and types. Values under the generic thermal sysfs interface are commonly represented in millidegrees Celsius; confirm the documented ABI for the target kernel and driver rather than guessing from the magnitude.

Trip points represent thresholds and their classes, such as active, passive, hot, or critical where supported by the platform. Cooling devices represent actuators, such as a processor cooling state or fan. A zone can link cooling devices to trips with weights. Therefore, a high cooling-device state is not by itself evidence of a sensor problem; it may be the expected response to a real temperature rise.

for zone in /sys/class/thermal/thermal_zone*; do
    [ -d "$zone" ] || continue
    printf '\n%s\n' "$zone"
    for attr in type temp mode policy; do
        [ -r "$zone/$attr" ] && printf '  %s=' "$attr" && cat "$zone/$attr"
    done
done

This is read-only enumeration. Keep the zone directory alongside each value in collected output so values from multiple zones are not merged accidentally. If an attribute is absent, record it as unsupported or unavailable rather than treating absence as zero.

Correlate cooling devices and governors

Inspect /sys/class/thermal/cooling_device* for each device’s type, cur_state, and max_state. The meanings of state numbers are device-specific. A larger state often means more cooling but cannot be assumed to have the same direction or magnitude for every driver. Some devices expose time-in-state and transition statistics that help establish whether the cooling policy changed during the incident.

The policy names describe which thermal governor controls a zone, and available_policies lists choices for that zone. Do not switch to user_space merely to silence a throttling symptom. Doing so may change the control authority and, on some platform designs, remove expected automatic mitigation. Read the platform driver and kernel ABI before writing any policy or trip point.

Collect measurements over time rather than sampling once. Correlate zone temperature, cooling-device state, CPU frequency or scheduler utilization, fan behavior, ambient conditions, and the workload’s throughput/latency. A sensor may report package temperature while another reports a board or skin zone. Map readings to the platform documentation before drawing conclusions.

hwmon overlap and sensor identity

The thermal subsystem may provide a hwmon interface for a thermal zone, and a hardware driver can expose related sensor data through its own hwmon device. Two files can therefore represent related readings with different names or update cadence. Record the name, label if present, units, and source path. Avoid averaging sensors that observe different physical areas or use different calibration.

The hwmon standard provides a common userspace-facing representation, but the driver determines which sensors and attributes are present. A missing temp*_input file can indicate the driver does not expose that channel, not a failed sensor. Check the kernel driver and hardware documentation for sensor mapping.

Distinguish thermal effects from other causes

CPU frequency can decline because of power limits, firmware policies, workload mix, scheduler placement, or an explicit userspace governor, even when thermal temperature is not near a displayed trip point. Conversely, a short thermal event can trigger a cooling state that remains active after the sampled temperature drops. Record frequency, utilization, power-related indicators, temperature, and cooling state on the same time basis.

Some CPUs expose package or core temperature through processor-specific drivers, while ACPI or firmware may publish additional platform zones. A temperature label does not necessarily map one-to-one to a physical sensor unless the driver or hardware documentation says so. Cross-check sensor identity with the vendor’s telemetry tools and use trends to find which zone changes with a controlled workload. Do not average package and chassis zones into one synthetic temperature.

Thermal pressure can affect scheduling and observed throughput beyond a simple frequency reading. A governor may reduce a cooling device state or influence CPU capacity, and firmware may independently enforce a power limit. If the kernel reports no thermal trip crossing but performance still falls, investigate power and scheduler signals alongside the thermal tree. Conversely, a thermal zone can be poorly calibrated or stale; a single extreme value should be verified before replacing cooling hardware.

For reproducible diagnosis, compare an idle baseline, a controlled workload, and the production workload. Ensure cooling pathways are unobstructed and firmware is current according to the platform vendor. Do not disable a fan or raise critical trip points to hide a symptom. Hardware protection is a safety mechanism, and changing it can risk data loss or equipment damage.

Reading trip points without rewriting them

Trip point attributes can be read where the driver exposes them. Capture type, temperature, hysteresis, and associated cooling devices. Hysteresis prevents rapid switching around a threshold, so a cooling action may remain active until the temperature falls sufficiently below the trip. A lack of a matching state transition at exactly the displayed threshold does not necessarily mean the driver is broken; the policy can include delay, trend, or other input.

The power allocator governor is a control loop with tunable parameters and requirements about passive trip points. Tuning it requires a platform-specific power model and controlled validation. Generic sysfs values are not a universal safe configuration. This article’s diagnostic snippets deliberately read the interface only; operational tuning belongs in the device vendor’s supported procedure.

Critical trips deserve special treatment. A driver may ask the system to shut down or reset when a critical temperature is reached. Do not test that behavior on a live production system by writing an emulated temperature or changing critical thresholds. If hardware testing is required, follow the vendor’s lab procedure, use a controlled test rig, and ensure an independent operator can restore the system. Emulated temperature can help exercise some software paths, but it does not validate the sensor, fan, or physical cooling response.

Test and document findings

Use a controlled, repeatable workload and collect aligned samples. Test fan or cooling response only using vendor-supported procedures and an appropriate lab environment. Compare the result with the system’s firmware, BIOS/UEFI, kernel, and driver versions. Include ambient conditions and whether the device is docked, charging, or in a different power profile.

An incident report should state which zone changed, its documented or inferred physical meaning, current temperature and trip thresholds, cooling-device states, observed performance impact, and any firmware or service events. Avoid presenting a zone index without the mapping. If values are implausible, compare against the platform’s own telemetry and investigate sensor calibration before changing policy.

Capture enough samples to establish sequence: workload starts, temperature rises, trip is crossed, cooling state changes, and performance responds. A policy may poll rather than react continuously, and hysteresis can create a gap between rising and falling transitions. Keep raw timestamps and sample intervals in the report. If a daemon manages power profiles, note its active profile and version because userspace policy can interact with kernel thermal decisions.

Linux thermal sysfs is a diagnostic model of temperatures and cooling relationships, not a universal hardware control panel. Identify each zone and actuator, measure transitions over time, correlate with workload evidence, and keep protective policy unchanged until the hardware-specific cause and supported remedy are established.

Related:

Sources:

Comments