Skip to content
FreeBSDDeep Dive Published Updated 7 min readViews unavailable

FreeBSD pmcstat: Profile CPU Hardware Events with hwpmc

Use FreeBSD hwpmc and pmcstat to count and sample CPU events, select supported counters, collect profiles, and interpret results without overclaiming.

FreeBSD’s hwpmc driver exposes processor performance-monitoring counters to userland. pmcstat can count or sample hardware events for the whole system or for selected processes. These counters help investigate questions such as whether a measured workload spends time retiring instructions, missing caches, or stalled on a particular CPU event. They are not universal performance metrics: event names and semantics vary by processor, kernel support, and counter implementation.

Hardware-counter profiling is most useful after a repeatable workload and a specific hypothesis exist. Start with elapsed time, throughput, latency distribution, CPU use, and system context. Then use a counter to distinguish competing explanations. A high cache-miss count alone does not prove that memory latency is the bottleneck; counts need normalization, workload boundaries, and a comparison against a controlled baseline.

Confirm that the system can measure

The hwpmc facility needs kernel support. The hwpmc(4) manual describes HWPMC_HOOKS and the hwpmc device as kernel configuration requirements; an installed system may provide the functionality in its active kernel or as a loadable component, depending on its build. Check the manual and kernel messages rather than assuming the module can be loaded on every architecture.

Record the CPU model, architecture, kernel version, and current kernel configuration before a measurement. Event sets differ across hardware generations, and a name copied from another vendor’s profiler may be unknown or may measure a different event. Some systems expose only a limited number of programmable counters. The tool can fail to allocate an event if the hardware or kernel cannot support the requested combination.

With the appropriate privileges and driver support, list events available on this machine:

sysctl hw.model
uname -a
pmcstat -L

Read the event descriptions in pmc(3) and the hardware-specific documentation for the processor. Use a name printed for the current host. Do not invent a generic event such as “cache misses” and assume its precise scope, unit, or privilege filtering is consistent across CPU models.

Choose counting or sampling

Counting answers how many times an event occurred over a measurement interval. Sampling periodically records where a selected event was observed, often with a call chain. Counting is useful for aggregate comparisons; sampling can locate hot regions but adds sampling overhead and statistical uncertainty. A short run or very low event rate can produce too few samples to support a detailed conclusion.

pmcstat uses lower-case and upper-case options for different modes. The manual defines -p and -s for process- and system-mode counting PMCs, and -P and -S for process- and system-mode sampling PMCs. A command placed at the end of the invocation can be measured as a target process. Begin with one event, one bounded workload, and a fresh output file.

A process-mode counting example uses a supported event name selected from the local list:

pmcstat -p EVENT_NAME_FROM_LOCAL_LIST ./benchmark 10000

Replace the event placeholder with a real event and the benchmark with a deterministic workload. Because this example’s command begins with a path rather than an option, pmcstat can distinguish the target from its own flags. If the installed release’s syntax differs, follow its local man page. Run the benchmark without counters first and record baseline timing, then repeat under the counter with identical input and environmental conditions.

For a sampling run, pmcstat can write an event log for offline analysis:

pmcstat -P EVENT_NAME_FROM_LOCAL_LIST -O /var/tmp/benchmark.pmclog ./benchmark 10000
pmcstat -R /var/tmp/benchmark.pmclog

The event name and workload are placeholders. Replace the event name with a value printed by pmcstat -L for the current processor. Protect the output directory from concurrent runs and preserve the exact executable and symbols. Sampling output is not a deterministic trace of every instruction; it is a statistical sample influenced by rate, interrupt handling, workload scheduling, and the hardware’s event semantics.

Keep measurements comparable

For a valid comparison, hold constant the executable build, input, CPU affinity, workload duration, background load, power policy, and kernel. Record CPU frequency and thermal behavior where relevant. FreeBSD power management or system load can change throughput during a run, so compare more than one repetition and report spread rather than selecting the fastest sample.

Process mode limits the measured scope to a target process and, depending on the option, possibly its descendants. System mode includes activity beyond the benchmark: interrupts, daemons, other users, and kernel work. System-wide counters can be useful to understand total CPU behavior but should not be attributed to one program without additional evidence.

Counters can overflow or be multiplexed depending on how many events are requested and what the hardware supports. Start with a single counter. If you need multiple events, run separate comparable trials or confirm from the installed manual and hardware documentation that the requested set can be scheduled simultaneously. A failed PMC allocation is a capability limitation, not a workload result.

Do not compare raw counts from different CPU models as if they were a stable unit. Even within one model, total event counts scale with runtime, core count, and instruction paths. Normalize only when the event semantics justify it, such as counts per operation or per elapsed interval. Publish the event name, mode, duration, sampling rate, processor, and pmcstat version with the result.

Read profiles as evidence, not proof

Sampling call chains are most useful when the executable and shared libraries have symbols. If the profile shows addresses without names, preserve the exact binary, debug symbols, and mapping information. A later rebuild can move code and make the old addresses misleading. The pmcstat manual supports offline processing from a log and can emit callchain or calltree data for compatible tools.

A high sample count in one function suggests that the sampled event frequently occurred while that code was executing. It does not show that the function caused every stall, nor that replacing it will improve end-to-end latency. Validate a proposed optimization with an A/B benchmark using the same workload and service-level metric. Check that the counter changes in the expected direction and that the user-visible metric also improves.

The human-readable output format is not promised stable by the manual. Avoid brittle scripts that parse column widths from console output. Store raw logs and use documented machine-readable or offline processing interfaces where available. Keep toolchain versions with the experiment so future analysis can reproduce symbolization and format interpretation.

Diagnose common failures

If pmcstat cannot allocate an event, check that the event name came from the local event list, the kernel supports hwpmc, and the selected mode exists on the CPU. A counter may be unavailable due to architecture, virtualization, privilege configuration, or another active measurement. Confirm the exact error before changing kernel settings.

If the profile is empty, verify the target command ran, the output path was writable, the event actually occurred, and the sampling period was suitable. A command that exits too quickly may finish before enough samples are collected. A system-mode run can collect data even when the target process is not as expected, so confirm whether the chosen mode measured the intended scope.

If results vary widely, quantify the variation. Check CPU frequency, interrupt activity, competing workloads, cache warmth, storage effects, and process placement. Use several repetitions and discard runs only for a documented reason. If a hardware event is unsupported or ambiguous, choose a better-defined event or rely on a different measurement method rather than forcing a conclusion.

A controlled profiling workflow

First write down the question and the metric that would confirm or falsify it. Capture a baseline without PMC instrumentation, then collect one event in process mode on the same workload. Repeat enough times to estimate noise. If aggregate counts identify a useful distinction, collect a sampling profile and resolve symbols using the matching build artifacts. Then change one code or system variable and rerun the same sequence.

A useful experiment record includes CPU model and architecture, FreeBSD release and kernel, event string, PMC mode, sampling parameters, exact command and input, executable hash, symbol set, wall-clock timing, system load, and all raw logs. This context is necessary because hardware counters are processor-specific and pmcstat output does not automatically explain the workload.

Performance acceptance should be based on the original user-visible metric, not on a more attractive counter value. A lower instruction count can accompany longer waits; fewer cache misses can accompany more instructions; more samples in a function can merely mean the process ran longer. pmcstat is a microscope for a defined hypothesis, not a universal quality score.

Related:

Sources:

Comments