Linux cgroup v2 memory.reclaim: Controlled Proactive Reclaim
Operate cgroup v2 memory.reclaim safely: understand its best-effort byte target, read memory events and PSI, and separate proactive reclaim from limits.
The cgroup v2 memory.reclaim file gives an administrator or delegated manager a way to request reclaim from a particular memory cgroup without waiting for a new allocation to cross a limit. It is useful for controlled experiments, workload-aware memory management, and testing a reclaim response. It is not a hard cap, a promise to free exactly the requested number of bytes, or evidence that the workload is under memory pressure. Those distinctions separate memory.reclaim from memory.high, memory.max, and pressure telemetry.
Find and verify the target cgroup first
Do not guess a cgroup path from a service name. Under systemd, ControlGroup reports the unit’s cgroup-relative path. Confirm that the cgroup v2 mount is where you expect, resolve the returned path, and make sure the unit is not the root cgroup. A service can have child cgroups, so decide whether the intended target is the unit’s parent or one specific child before issuing any write.
unit='example.service'
cg=$(systemctl show --property=ControlGroup --value "$unit")
mount='/sys/fs/cgroup'
case "$cg" in
''|'/')
printf 'No non-root cgroup path for %s: %s\n' "$unit" "$cg" >&2
exit 1
;;
/*) ;;
*)
printf 'Unexpected cgroup path: %s\n' "$cg" >&2
exit 1
;;
esac
dir="${mount}${cg}"
test -d "$dir" || { printf 'Missing cgroup: %s\n' "$dir" >&2; exit 1; }
test -r "$dir/memory.current" || { printf 'No memory controller view\n' >&2; exit 1; }
test -e "$dir/memory.reclaim" || { printf 'memory.reclaim is unavailable\n' >&2; exit 1; }
printf 'Target: %s\n' "$dir"
cat "$dir/memory.current"
cat "$dir/memory.stat"
cat "$dir/memory.events.local"
if test -r "$dir/memory.pressure"; then
cat "$dir/memory.pressure"
else
printf 'memory.pressure is unavailable in this kernel or cgroup view\n' >&2
fi
This inventory does not trigger reclaim. It confirms the chosen cgroup and captures useful context first. The path /sys/fs/cgroup is common but not universal; use the actual cgroup v2 mount on the host. Containers may see a namespaced or delegated subtree rather than the host’s full hierarchy. memory.events is hierarchical by default, while memory.events.local isolates events in the selected cgroup, which is often easier to correlate with one experiment. The memory.stat keys are not a fixed-position table: parse by key, and tolerate kernel versions that add fields.
What a reclaim write means
memory.reclaim is a write-only, nested-keyed cgroup v2 interface. A byte quantity requests that the kernel try to reclaim that amount from the target cgroup:
printf '%s\n' '256M' | sudo tee "$dir/memory.reclaim" >/dev/null
Choose a small amount on an isolated canary first. Capture the before-state, make one request, then read memory.current, memory.stat, memory.events.local, and memory.pressure again while observing application latency and throughput. Keep the output and time window together so a decrease in usage is not attributed to reclaim if the workload simply exited, freed its own allocations, or entered a different phase.
The requested amount is not an exact postcondition. The kernel may reclaim more or less than the request; when it reclaims fewer bytes than requested, the write can fail with EAGAIN. A successful write therefore means the request completed according to the interface, not that a precise number of resident bytes vanished or that application performance was unchanged. Some pages are not immediately reclaimable, and reclaim may change the balance between anonymous memory, file cache, swap, and other accounted components. Inspect the relevant memory.stat keys and swap controls instead of treating memory.current as an explanation by itself.
This proactive reclaim path also does not represent memory pressure in the cgroup. In particular, reclaim caused by this interface normally does not trigger the socket-memory balancing behavior that is associated with pressure-driven reclaim. For network services, that difference can make an artificial test unlike real pressure and can hide a behavior that appears during actual memory contention.
Keep reclaim separate from limits and protection
The memory controller exposes different controls for different contracts:
| Interface | Operational meaning | What it does not promise |
|---|---|---|
memory.reclaim |
Request proactive reclaim from a cgroup now | Exact bytes reclaimed, lasting low usage, or pressure detection |
memory.high |
Throttle the cgroup and put it under reclaim pressure above a boundary | A hard cap or OOM protection under every extreme condition |
memory.max |
Enforce a hard usage limit; reclaim is attempted and a cgroup OOM can follow if usage cannot be reduced | Graceful application behavior when the limit is reached |
memory.low / memory.min |
Best-effort or hard reclaim protection, subject to ancestor and sibling competition | Unconditional protection when configured values are overcommitted |
Use memory.high when a manager needs an ongoing pressure boundary that it can monitor and adjust. Use memory.max when containment is the requirement and the failure mode has been designed and tested. Use memory.reclaim for a one-off or policy-driven reclaim request, not as a substitute for either limit. A workload can show a large memory.current value while operating efficiently; memory usage alone is not a measure of user-visible pressure.
Build a measurement loop, not a one-shot command
Before an experiment, record the target cgroup, kernel release, workload version, request size, relevant ancestor limits and protections, memory.swap.current and memory.swap.max where present, and the workload’s latency/throughput baseline. Take counter snapshots rather than comparing only absolute totals. For event files, preserve both the initial and final values and compute deltas; event counters are cumulative and some are hierarchical.
After the request, check whether the application changed behavior. Read cgroup-local and descendant event counters, pressure-stall information, reclaim/writeback activity, swap usage, and storage latency. memory.pressure and system-wide /proc/pressure/memory answer a different question from memory.current: they describe time lost to memory contention, not simply allocated bytes. If PSI rises while the request is applied, that is evidence of a workload impact to investigate, not by itself proof that the request caused it. Use aligned time windows and compare against a control run.
Repeat the experiment with the same workload phase and a conservative request size. If results vary, inspect working-set changes, dirty file cache, pinned or otherwise unreclaimable pages, ancestor constraints, and concurrent sibling activity. Reclaim protections such as memory.low and memory.min affect which pages are eligible in context; they are hierarchical, and overcommitment changes how protection is shared. Avoid weakening protections or changing multiple controls in the same experiment just to force a desired outcome.
Production rollout and rollback
If an automated memory manager will write memory.reclaim, first validate that the selected cgroup belongs to the intended workload and that the manager’s authority is limited to that subtree. Bound each request, rate-limit repeated requests, and record the requested bytes, result, errno, before/after counters, and workload SLOs. A manager that retries every EAGAIN at full size can create a feedback loop that increases pressure or wastes CPU. Use explicit retry limits and backoff, and stop when the workload is degraded or the underlying memory composition is not reclaimable.
Start with a canary and an explicit kill switch. Roll back the automation by disabling its writer or policy; a reclaim operation already in progress cannot be undone by restoring a previous memory.current value. Do not treat a low post-reclaim working set as success unless the same traffic, latency, throughput, error rate, and recovery behavior remain acceptable. For a persistent policy, retain enough telemetry to distinguish a deliberate reclaim experiment from natural memory pressure and memory-limit enforcement.
memory.reclaim is most valuable when it closes a measured feedback loop: identify a cgroup, make a bounded request, observe the kernel and application, and keep or revert the policy based on repeatable evidence. The interface enables experimentation; it does not remove the need to understand the workload’s working set or failure modes.
Related:
- Linux DAMON: Measure Memory Access Patterns Before Tuning
- Linux Pressure Stall Information: Measuring CPU, Memory, and I/O Contention Directly
Sources: