Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux perf c2c: Find Cache-Line Contention Before Blaming the Lock

Use Linux perf c2c to locate cache-to-cache traffic, validate false-sharing hypotheses, and separate hardware sampling limits from application contention.

High CPU utilization does not always mean a thread is doing useful computation. On multicore systems, CPUs can spend time transferring cache lines that contain data written by other cores. perf c2c is a Linux performance tool for analyzing cache-to-cache traffic and HITM events, helping locate cache lines with high contention. It can support a false-sharing diagnosis, but it does not automatically prove that adjacent variables are the root cause or that padding a structure will improve the application.

The tool relies on processor performance-monitoring capabilities, kernel and userspace support, and sufficient symbol/debug information. Its event coverage and sampling behavior vary by architecture and CPU generation. Treat a c2c report as a hypothesis generator that must be correlated with source layout, workload phases, and application-level performance.

What cache-to-cache evidence means

Modern processors maintain coherent caches. If two CPUs repeatedly access a line and at least one modifies data in it, ownership can move between caches. This is normal for shared mutable data, locks, reference counts, and per-core counters that happen to occupy the same cache line. The performance problem is not sharing by itself; it is costly ownership movement on a hot path.

False sharing occurs when independently used variables happen to occupy the same cache line. Threads may never logically communicate through those variables, yet writes to one field invalidate another core’s copy of the line. True sharing occurs when threads intentionally access the same field or synchronization variable. Both can generate cache-to-cache traffic. Fixes differ: false sharing may respond to layout or sharding, while true sharing may require algorithmic reduction of shared writes.

HITM stands for a cache-to-cache hit in modified state. A high HITM count can indicate line ownership traffic, but its exact meaning depends on the PMU event and architecture. It does not tell you whether a C struct field is incorrectly placed without decoding the address and access offsets. Sampling is statistical, and some CPUs require specific load-latency or precise-store facilities. A report with zero samples can mean missing hardware support, permissions, event constraints, or a workload that did not exercise the condition.

Prepare a useful recording

Start with a stable, representative workload. Record the CPU model, kernel, perf version, CPU topology, process affinity, compiler build, and whether symbols are available. A short recording during idle startup is not comparable to a saturated service interval. Keep the duration bounded because sampling adds overhead.

The kernel false-sharing guide demonstrates a basic workflow:

perf c2c record -ag -- sleep 10
perf c2c report --call-graph none -k vmlinux

Replace sleep 10 with the target command or a carefully selected workload interval when appropriate. The -k vmlinux option helps resolve kernel addresses when matching symbols are installed; userspace symbols depend on the executable and debug information. Do not assume a report can decode source lines from a stripped binary.

Check whether perf c2c record -e list reports supported events on the target CPU. Hardware support differs: Intel, AMD, Arm, and PowerPC use different event facilities and may have different constraints. On some architectures, sampling is explicitly statistical and cannot capture every memory operation. Respect perf permissions and production policy; changing perf_event_paranoid globally is not an acceptable first diagnostic step.

Capture a baseline run and a comparison run with the same workload, CPU affinity, runtime, and load. If the optimization under investigation changes code layout, rebuilds, or thread placement, record those differences. Changes in cache alignment or binary address can change sampling attribution and invalidate a naïve comparison.

Read the report in layers

Begin with the cache-line table. Identify lines with high HITM or other contention-related samples, then examine the access offsets and participating symbols. Compare the line address and offsets with the structure layout using debug information and tools such as pahole. For a global symbol, consult matching System.map or debug symbols. Ensure these artifacts correspond to the exact kernel or executable that was sampled.

Next inspect the load/store views and call graphs. A cache line accessed from multiple code paths can be shared by unrelated subsystems. Determine which fields reside at the reported offsets and which threads or CPUs touch them. The kernel guide gives examples where several fields in a large structure share a line; a hot counter can therefore interfere with a seemingly unrelated pointer or state field.

The report’s top row is not automatically the best fix target. Consider sample count, percentage, instruction context, and the workload’s service objective. A line with a large share of HITM but negligible absolute activity may not explain wall-clock latency. Conversely, a modest contention count on a critical lock can affect tail latency more than a larger count on a background task.

Validate false-sharing hypotheses

For each candidate line, map the sampled byte offset to a specific field and reason about writers. If two fields are independently updated on different CPUs but share a line, the candidate is plausible false sharing. If both CPUs update the same field or lock, it is true sharing. If one writer is a kernel worker and another is an application thread, verify that both addresses and symbols refer to the same physical line and that sampling attribution is correct.

Test a change in an isolated benchmark. Options include per-thread counters with aggregation, separating hot mutable fields, using cache-line alignment where the architecture and ABI permit, reducing write frequency, or changing the synchronization algorithm. Do not add padding indiscriminately: it increases memory footprint, can hurt cache locality, and may change public structure layout or ABI.

Use repeated runs and compare throughput, tail latency, CPU cycles, instructions, cache misses, and c2c samples. A lower HITM count without better application performance may mean the line was not the dominant bottleneck. A better latency result with similar samples may mean the optimization helped another part of the workload. Keep both hardware evidence and service-level outcomes.

Common interpretation errors

Calling every HITM line false sharing. HITM is evidence of modified-line traffic, not a source-level diagnosis. Determine whether fields are independent or intentionally shared.

Comparing samples across different CPUs as absolute counts. PMU facilities and sampling coverage vary. Use same-host before/after measurements and report the processor model.

Ignoring symbol mismatch. A wrong vmlinux, stripped executable, or stale debug package can assign an address to the wrong source line. Confirm build IDs and exact binaries.

Treating sampling percentages as wall-clock percentages. The tool samples hardware events; reported shares are not automatically CPU utilization or request latency.

Padding as the default remedy. It may reduce false sharing but also increase footprint and move contention elsewhere. Make a narrow change and benchmark it.

Running production profiling without a bound. Profiling has overhead and can collect sensitive addresses or call paths. Scope recording, protect outputs, and follow operational policy.

Acceptance and reporting

A defensible c2c report should include CPU model and architecture, kernel/perf versions, recording command and duration, workload identity, event support, symbol provenance, top candidate cache lines, decoded field offsets, and before/after application metrics. State whether the diagnosis is false sharing, true sharing, or unresolved cache-line contention. If the PMU does not support required events, stop rather than presenting an empty report as a clean bill of health.

For an accepted optimization, preserve the test harness and build artifacts, repeat measurements under representative concurrency, and confirm no ABI or memory-footprint regression. Re-profile after major compiler, kernel, or CPU changes because cache layout and PMU behavior can change.

The useful role of perf c2c is precise localization. It narrows the search from “the CPU is slow” to “these cache lines are moving between these execution contexts.” Source-level reasoning and controlled measurements determine whether that movement is accidental, necessary, or worth fixing.

Related:

Sources:

Comments