Linux DAMON: Measure Memory Access Patterns Before Tuning
Use Linux DAMON and damo to inspect approximate memory hotness, tune bounded sampling, and validate reclaim ideas without confusing telemetry with guarantees.
Linux DAMON, the Data Access MONitor, estimates how memory regions are accessed over time. It is useful when a service’s resident set is large but operators do not know which parts are active, when a working set changes, or whether a memory-management policy is worth testing. DAMON can also drive kernel-side actions through DAMOS, its operation-scheme layer. Treat those as separate stages: first measure and validate the workload, then consider bounded actions. A heat map is not proof that reclaiming a region will be harmless.
What DAMON measures, and what it does not
DAMON uses an operations set suited to the address space being observed. The vaddr set monitors a process’s virtual address space, fvaddr monitors configured virtual ranges, and paddr monitors physical memory. A context holds a monitoring request and its schemes; a kernel thread called kdamond executes that context. The choice matters: process-level questions usually call for vaddr, while system-wide questions may call for paddr and require different permissions and interpretation.
Its core design is adaptive region-based sampling. Instead of checking every page on every interval, DAMON samples one page in each region and treats adjacent pages as having a similar access pattern. It merges and splits regions as observations change, within configured minimum and maximum region counts. This bounds overhead, but it also makes the result an estimate: an unrepresentative sample or a rapidly changing pattern can hide detail. More regions can improve resolution while raising monitoring cost; fewer regions reduce overhead while smoothing differences.
The reported access count and age answer different questions. Access frequency summarizes observations during an aggregation interval. Age describes how long a region’s size and access frequency have remained sufficiently similar under DAMON’s change rules; it is not the age of the allocation or the process. Neither value is a CPU cache-miss counter, a precise per-page access trace, or a direct measure of application latency. Correlate snapshots with request latency, throughput, PSI, working-set changes, and workload phase boundaries.
Check support before collecting data
Kernel configuration, operation-set support, permissions, distribution backports, and the installed damo version all affect what is available. Start with the capability report rather than assuming that a command from another host will work:
sudo damo report sysinfo
Confirm that the report lists the interface needed for monitoring, and inspect the installed damo help and version if an option differs from the current upstream guide. DAMON can be present in a kernel while a particular interface or operation is unavailable. Access checks may use page-table accessed bits; the kernel documentation notes possible interaction with idle-page tracking, so include that caveat in a system-level measurement plan.
For a controlled first capture, identify one representative, long-running workload process. For a systemd service, MainPID can be a useful starting point, but it does not necessarily represent every worker in a multi-process service. Do not pass an empty, zero, or ambiguous PID to a privileged collector:
service_name='example.service'
pid=$(systemctl show --property=MainPID --value "$service_name")
case "$pid" in
''|0|*[!0-9]*)
printf 'No single usable MainPID for %s: %s\n' "$service_name" "$pid" >&2
exit 1
;;
esac
printf 'Candidate target: %s (PID %s)\n' "$service_name" "$pid"
This only selects a candidate. Verify that it is the process whose mappings and workload phase you intend to study. A short-lived PID can exit or be reused; a process can also create or remove mappings while the capture is running. For a service pool, decide whether to monitor a representative worker, each worker in a controlled experiment, or a system-level address space. Do not silently combine unlike workloads into one conclusion.
Capture a baseline before enabling any memory action
The upstream damo utility can record an already-running process by PID. Full recording requires root and uses perf or trace-cmd to collect DAMON access results; first confirm the prerequisites on the target. Output files are root-owned with restrictive permissions by default, which is appropriate because the records can disclose process address ranges and memory behavior. Store them with the same care as diagnostic dumps.
sudo damo report sysinfo
sudo damo record "$pid"
Allow the capture to span representative steady-state and relevant burst or idle phases, then stop it deliberately with the foreground command’s interrupt handling. Keep the workload and DAMON configuration notes with the record: kernel build, damo version, PID/service identity, start and stop times, workload phase, and any CPU or memory limits. Without these, snapshots from two runs may not be comparable. Avoid leaving a broad or indefinite capture running by accident; long recordings can produce large files, so use the installed version’s duration, snapshot, or output-flush options when appropriate.
Inspect the record rather than jumping to a reclaim policy:
sudo damo report access --input damon.data
sudo damo report damon --input_file damon.data
The access report exposes sampled regions and their access behavior; the status report helps establish which configuration produced them. Compare multiple snapshots. A single observation during startup, garbage collection, cache warm-up, a backup, or a traffic spike is not a stable working-set profile. If the same service has multiple materially different phases, label and compare those phases rather than averaging away the distinction.
Tune accuracy against overhead and workload timescale
The aggregation interval determines how long DAMON accumulates samples before reporting a snapshot. It should match the time scale of the question. If a workload’s hot set changes every few seconds, a very long aggregation interval can blur that transition; an interval that is too short may capture too few accesses to distinguish regions. Sampling interval controls temporal resolution and contributes to monitoring cost. Region-count bounds control spatial granularity and overhead. There is no universal setting that makes all workloads both precise and free.
Begin with defaults or a conservative, documented configuration and compare repeated captures. Change one parameter at a time. Check whether the conclusions survive a modest interval change, a second workload run, and a different phase. If a region is called cold only under one short snapshot, regard it as a hypothesis. DAMON’s region model deliberately trades fine-grained certainty for bounded monitoring work.
Virtual-address monitoring also has mapping boundaries. DAMON updates its view of a process’s address space periodically rather than continuously. Dynamic mappings, allocator arenas, memory-mapped files, and process restarts can all affect what is observed. Preserve mapping and process identity context when investigating a surprising region. For longer-term monitoring, ensure that collection cadence and file rotation are sized for the desired retention instead of treating raw records as free telemetry.
Treat DAMOS as a separate, guarded change
DAMOS expresses a high-level access pattern and a requested action. Supported actions depend on the operations set. Some actions can advise or reclaim memory, while stat collects statistics without changing the target pages. A scheme may support filters, quotas, and watermarks; these controls can bound how much action is attempted or when a scheme runs, but they do not prove that the action is safe for a particular application.
Use a staged rollout:
- Record an observational baseline and define a workload-level success metric, such as tail latency, throughput, or memory pressure.
- Validate that the candidate access pattern recurs across representative runs and is not a transient startup or maintenance phase.
- Start with statistics-only observation where supported; review tried-region and scheme statistics before considering a mutating action.
- If testing an action, use a canary workload, narrow access pattern, conservative quota, and an explicit rollback procedure. Watch application SLOs, PSI, reclaim activity, swap behavior, and OOM events together.
- Stop or revert the experiment if the workload regresses, monitoring results drift, or the scheme’s applied regions differ from the intended target.
For example, DAMON-based reclaim is an optional kernel feature with its own parameters and behavior; it should not be inferred from a generic DAMON access report. Likewise, a pageout or cold scheme does not guarantee a particular RSS reduction, swap destination, or latency outcome. Reclaimability depends on page type, references, memory pressure, kernel configuration, and concurrent activity. Benchmark and validate each operational policy on the actual kernel and workload before making it persistent.
Interpret results with independent signals
Use DAMON to answer “which regions appear active, and for how long?” Use PSI and application telemetry to answer whether the system or service is experiencing contention or user-visible impact. Use process maps, cgroup accounting, reclaim counters, and storage latency to explain why a memory policy may help or hurt. If you need to tell a cooperating process what to do with ranges, process_madvise() is a separate interface with its own permissions and advice semantics; DAMON does not turn its approximate observations into authoritative application intent.
The production-grade outcome is not the most aggressive scheme. It is a reproducible observation, a clearly stated uncertainty, an action bounded by quotas and rollback, and evidence that the measured workload improved without unacceptable regressions.
Related:
- Linux process_madvise(): Advising the Kernel About Another Process’s Memory
- Linux Pressure Stall Information: Measuring CPU, Memory, and I/O Contention Directly
Sources: