Skip to content
FreeBSDDeep Dive Published Updated 7 min readViews unavailable

FreeBSD CPU Temperature Monitoring: Use coretemp and amdtemp Carefully

Monitor FreeBSD Intel and AMD CPU temperature sensors, load supported drivers, interpret sysctl values, and avoid false thermal alarms.

FreeBSD can expose on-die CPU temperature readings through hardware drivers such as coretemp(4) and amdtemp(4). These readings are useful for trend monitoring and hardware diagnosis, but they are not a universal chassis-temperature interface. Driver support depends on processor family, firmware, and the running kernel. A missing sensor does not prove that a CPU is cool, and a displayed value is not a substitute for the system’s own thermal protection.

Treat temperature telemetry as a measurement channel with explicit coverage. Record which CPUs expose a sensor, when the reading is sampled, and which units the driver reports. Do not copy thresholds from another processor model or assume that every core number maps to a physical socket or core layout in an obvious way.

Load only the driver supported by the CPU

The FreeBSD manual documents coretemp for supported Intel Core and newer processors and amdtemp for supported AMD processors. Inspect the CPU model and loaded modules before changing startup configuration:

sysctl hw.model
kldstat
dmesg | grep -i -E 'coretemp|amdtemp|temperature|cpu'

If the driver module is available and appropriate for the installed hardware, load it for a bounded test:

kldload coretemp
sysctl -a | grep -E '^dev[.]cpu[.].*temperature'

Use kldload amdtemp instead for a supported AMD CPU. Do not load both blindly on an unknown platform, and do not assume a successful kldload means the hardware exposed usable sensors. Confirm the kernel message and resulting sysctl nodes. Some drivers may be built into the kernel and therefore not appear as a separately loaded module.

For a persistent module, verify the exact driver name and local loader.conf(5) guidance, then apply one deliberate configuration change. A line such as coretemp_load="YES" is appropriate only when the module is present and supported. Keep a rollback path if a custom kernel or remote system fails to boot after a loader configuration change.

Read sensor values without inventing topology

The coretemp(4) manual documents readings beneath dev.cpu.%d.temperature. Query the actual nodes reported by the running host:

sysctl dev.cpu.0.temperature
sysctl -a | grep -E '^dev[.]cpu[.][0-9]+[.]temperature'
sysctl kern.smp.cpus kern.sched.topology_spec

The value often includes a temperature unit in its printed representation, but parsers should inspect the exact output on the target release rather than stripping text based on an assumption. The index %d is a FreeBSD CPU device index, not a portable promise about socket order, firmware order, or physical placement. kern.smp.cpus records the CPU count; it is not a socket/core map. Where available, inspect kern.sched.topology_spec and corroborate the result with the platform documentation before assigning physical labels.

AMD readings need an additional driver-specific check. The cited FreeBSD 15.1 amdtemp(4) manual lists support for AMD families 0Fh, 10h, 11h, 12h, 14h, 15h, 16h, and 17h; it describes per-core readings for family 0Fh and a shared package sensor for families 10h through 17h. The manual also warns that family 10h and later values are a non-physical relative scale, not an actual die or case temperature. Do not label those values as physical Celsius measurements or apply ordinary temperature thresholds without validating the processor-specific meaning. Check the manual for the installed release and CPU family rather than extrapolating this support list to newer AMD families.

Do not infer an absent core node’s value from a neighboring core. A CPU may be offline, unsupported, or not represented by that driver. A machine may expose a sensor on some CPUs but not others. Record missing readings as missing telemetry, not zero degrees or healthy status.

Build trend monitoring, not a one-number alarm

Take repeated samples under idle and representative load. A single sample cannot establish a fan curve, thermal throttling event, or sensor failure. Pair readings with clock/frequency observations, workload, ambient conditions, fan or chassis sensors if available, and system logs. Sensor response time and sampling frequency can smooth or miss short-lived peaks.

Temperature thresholds are hardware-specific. Use the processor vendor’s supported thermal specifications and platform documentation to set warning and action levels. Do not use an arbitrary generic Celsius threshold as an emergency shutdown trigger. The kernel and processor may implement throttling or shutdown independently; external monitoring should raise an alert early enough for an operator to investigate, not race the hardware’s safety mechanisms.

For a periodic collection script, bound the command and preserve both the raw value and timestamp. Validate that the node exists before parsing it and distinguish command failure from a legitimate numeric reading. An alert pipeline should report sensor absence, stale samples, and out-of-range values separately. If the management agent restarts, verify that the graph does not silently treat a gap as normal temperature.

Diagnose absent, implausible, or unstable readings

When no nodes appear, verify the CPU model, kernel driver availability, module load result, and dmesg output. Check the installed coretemp(4) or amdtemp(4) manual for processor support. Do not download an untrusted kernel module or substitute a different vendor driver because a monitoring guide names a similar CPU family.

If values are implausible, compare a cold-start idle sample with a controlled load and use a second hardware telemetry source when available. Do not immediately conclude the sensor is broken: virtualization can expose a different CPU model, firmware can mask hardware features, and the host may not provide guest-visible thermal registers. A VM’s temperature reading is not necessarily the physical host CPU temperature.

If only one CPU node changes, verify topology and driver semantics before treating the discrepancy as a failed core. Look at sustained behavior and corroborating performance symptoms such as frequency changes or thermal warnings. Avoid stress-testing production hardware solely to force a high reading. Use a bounded workload in a maintenance window with temperature and shutdown behavior monitored.

Account for virtualization and operational ownership

Guests commonly cannot read host sensors directly. A hypervisor may expose virtual CPU features while deliberately withholding thermal information. If a guest has no readings, collect temperature at the host’s supported management layer instead of trying to load arbitrary guest drivers. Document where each metric originates so dashboards do not compare host and guest data as equivalent.

On bare metal, sensor visibility can change after BIOS updates, kernel updates, CPU replacement, or a hardware revision. Include the driver, CPU model, FreeBSD release, sysctl OID, sample cadence, and unit in monitoring metadata. Revalidate after a kernel upgrade and alert on a disappearing node. Do not assume configuration that worked on one motherboard will work on another system with the same CPU family.

Keep a baseline collected when the machine is known to be healthy. Record idle readings after a stable warm-up, then compare them with controlled workload samples under similar ambient conditions. This baseline is for detecting change on the same host, not for comparing unrelated processor models. A newer BIOS, different fan curve, or replacement heatsink can legitimately shift readings, so annotate hardware and firmware changes alongside the graph.

Monitoring should detect stale data as explicitly as high data. If the collector stops, the last successful temperature can remain on a dashboard and look normal. Attach a sample timestamp, enforce a maximum age, and alert when the sysctl query fails or returns an unparseable value. Recovery should restore data collection and confirm a new timestamp, not merely clear the alarm.

Use least privilege for collectors. Reading a sysctl node is generally less intrusive than changing a tunable, but collection agents should not receive broad root access merely to sample a temperature. Use the platform’s supported service account and permission model, and verify the final metric from the same account used in production.

Acceptance checks

An operational acceptance test should show the expected driver loaded or built in, list the available CPU temperature OIDs, return parseable readings with a timestamp, and produce an explicit missing-sensor alert when the source is unavailable. Compare a controlled idle and load window, verify that dashboards label units and host identity, and confirm the monitor does not initiate unsupported power-management actions.

These drivers expose hardware measurements; they do not guarantee that all thermal risks are covered. A robust system correlates CPU readings with platform sensors, workload, and vendor guidance, while leaving thermal control to the hardware and operating system mechanisms designed for it.

Related:

Sources:

Comments