Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux perf Event Counting: PMUs, Multiplexing, and Measurement Error

Measure Linux hardware and software events with perf while accounting for PMU scope, counter multiplexing, lost samples, permissions, and noise.

Linux perf exposes hardware performance-monitoring units (PMUs), software events, tracepoints, and sampling facilities through a common interface. A counter such as cycles or instructions can answer a focused question, but a counter value is not a self-explanatory performance diagnosis. Event definitions vary by processor, the number of programmable counters is limited, multiplexing can scale estimates, and the observer itself can perturb a short workload.

A trustworthy measurement starts with scope and event semantics. Are you counting one process, its descendants, one CPU, or the whole system? Is the event a generalized hardware event, a model-specific event, or a software counter? Did it run for the full interval, or was it multiplexed with another event? Did the workload and CPU frequency state remain comparable? Without these facts, an impressive-looking number can be numerically precise and scientifically misleading.

Counting and sampling answer different questions

Counting accumulates event totals and returns them through reads. Sampling periodically records a sample when a counter reaches a configured period or another sample condition. perf stat is usually the first tool for aggregate counts; perf record and perf report are used to attribute sampled activity to code locations when symbols and call chains are available.

Counting has low data volume but no direct timeline. A total of 10 billion cycles over an interval does not reveal whether the work was evenly distributed or concentrated in one stall. Sampling can locate hot code but introduces buffer management, sample loss, and additional overhead. Do not infer that a high sample count equals an exact count of every event; samples are statistical observations unless configured and interpreted as a precise event stream.

Generalized events make common requests more portable, but the mapping can differ across architectures and processor models. cycles, instructions, cache-misses, and branch events may be backed by different hardware definitions, constraints, or errata. Check perf list, the CPU vendor’s PMU documentation, and the perf_event_open(2) manual before comparing results across machines.

Event scope and grouping

perf_event_open() accepts a process ID and CPU selector that determine where an event is measured. A task-scoped event follows a task, optionally including inherited children; a CPU-scoped event observes activity on a CPU and may include unrelated work unless filters exclude it. Cgroup monitoring and system-wide collection have their own semantics and permissions. State the scope in every benchmark report.

Events can be grouped to request that they be scheduled together. This matters when ratios require simultaneous observations, such as instructions and cycles for an interval. A group may fail to schedule if the PMU cannot fit its members or if constraints conflict. Splitting events into separate runs avoids multiplexing but changes execution conditions; grouping avoids some temporal mismatch but does not create additional hardware counters.

perf list | rg -i 'cycles|instructions|cache-misses'
perf stat -r 5 -e task-clock,context-switches,cycles,instructions -- sleep 10

The first command discovers event names available to the installed tool. The second is a short system smoke test, not a workload benchmark; permissions or virtualized PMUs may deny hardware counters. For a real comparison, replace sleep 10 with a deterministic program, repeat trials, pin or control the workload only if that reflects the production question, and capture the exact command and CPU model.

Multiplexing and scaled estimates

When more events are requested than can be counted simultaneously, the kernel can multiplex them over time. perf_event_open() exposes time_enabled and time_running values when the requested read format includes them. Tools commonly scale a raw count by the ratio of enabled time to running time. This estimates what the total might have been had the event run for the entire interval, assuming the sampled periods are representative.

That assumption can fail when phases differ. If an event runs mostly during a quiet phase and is disabled during a burst, scaled totals can be biased. Very low running time makes the estimate noisy. A percentage indicating that an event was multiplexed is part of the result, not harmless output decoration. For ratios such as IPC, simultaneous grouping is preferable when the PMU can schedule the group; otherwise report the limitation and repeat measurements in controlled separate runs.

Do not compare a scaled count from one CPU with an unscaled count from another. Check output for <not counted>, <not supported>, multiplex percentages, and warnings. Event constraints can arise from counter slots, fixed versus programmable counters, precise-event requirements, privilege filters, or hybrid core types. On a heterogeneous processor, a “CPU” event may not have identical meaning on performance and efficiency cores.

Sampling buffers and lost data

In sampling mode, perf can map a ring buffer and write sample records that userspace later consumes. The metadata page communicates ring-buffer positions and event attributes. The consumer must keep up with the producer; if the buffer fills, samples can be lost. The report should preserve lost-sample information and the sample period, frequency mode, call-chain configuration, and filters.

Call chains depend on unwind information, frame pointers, architecture support, and kernel configuration. A missing stack is not automatically evidence that the code ran without callers. JITs need symbol registration or other support if samples should resolve to generated code. Kernel addresses may be restricted, and software versions can differ in symbol formats.

perf record itself consumes CPU and memory, and a high sampling frequency can distort a latency-sensitive workload. Begin with a conservative sample rate and bounded duration. Compare a run with and without profiling to estimate observer effect. For workloads with strict latency budgets, use controlled replicas or hardware tracing designed for the platform rather than collecting an unlimited sample stream from a live service.

A disciplined benchmark loop

First define the hypothesis: for example, “the new parser uses fewer instructions per input record without increasing branch misses or p99 latency.” Choose events that measure that hypothesis, not every event in perf list. Make the input, warm-up, CPU placement, governor state, kernel version, and compiler build reproducible. Record whether background load was present and whether the process was inside a container or cgroup.

Run multiple independent trials and report median plus spread. Avoid subtracting two large noisy totals to claim a tiny improvement. Keep wall-clock time, CPU time, throughput, and application latency alongside PMU counters. A reduction in cycles may mean less work, a lower frequency, fewer completed records, or a different execution path; normalize by a meaningful unit such as cycles per request and verify work completed.

Use perf stat -r for repeated measurements and inspect enabled/running percentages. Use perf record only after aggregate counts identify a question about where activity occurs. Then rerun with targeted sampling, preserve the perf.data file and command line, and inspect lost samples. Confirm any proposed optimization with an end-to-end benchmark rather than optimizing a sampled leaf function in isolation.

Permissions and platform constraints

The perf_event_paranoid setting and capabilities can restrict which events a user may open. A denied event is not fixed by changing the event name repeatedly. Ask the system administrator for the narrow permission needed, or use permitted software events and process-scoped measurements. Do not recommend lowering a host-wide sysctl as a routine debugging step; access policy belongs to the platform owner.

Virtual machines may virtualize, filter, or omit PMU features. Containers can have different visibility from the host, and cgroup permission does not necessarily permit host-wide collection. Record the hypervisor and kernel configuration. An event shown by the tool can still be unsupported on the current CPU, and raw encodings must never be copied blindly between processor models.

Acceptance criteria for a measurement

For every reported result, preserve kernel and perf versions, CPU model and topology, event names and modifiers, command scope, trial duration, repeat count, enabled/running ratios, sampling loss, and workload throughput. State whether values are raw or scaled. Re-run a smaller event set if multiplexing is material, and validate conclusions with application-level latency or throughput.

A counter experiment passes only when the workload is equivalent, measurements are repeatable, event semantics are understood for the processor, and observed changes support the original hypothesis. perf provides a powerful measurement interface; it does not guarantee that an event means the same thing on every CPU or that a correlation identifies a root cause. Keep the methodology with the code change so future kernel and hardware upgrades can repeat the experiment.

Related:

Sources:

Comments