Linux Multi-Gen LRU: Generations, Reclaim, and Evidence-Based Tuning
Understand Linux Multi-Gen LRU aging, eviction, runtime controls, and memory-pressure evidence before changing reclaim policy on production hosts.
Linux page reclaim must decide which cached or anonymous pages are least valuable when memory becomes scarce. Multi-Gen LRU, usually abbreviated MGLRU, changes how the kernel groups pages by recency so reclaim can reason about working sets with fewer list operations than the traditional active/inactive LRU scheme. It is a reclaim implementation, not a memory limit, a promise that cold pages remain resident, or an application-level cache. Results depend on the workload, kernel configuration, architecture, memory cgroups, swap policy, and the pressure pattern that caused reclaim.
The operational question is not whether MGLRU is intrinsically faster. It is whether the running kernel supports and enables it, whether reclaim is affecting the workload, and whether a controlled comparison shows better latency or memory efficiency without unacceptable eviction or CPU cost.
What a generation represents
The kernel observes page-table access information and organizes eligible pages into generations that approximate recency. Newer generations contain pages observed more recently; older generations are candidates for reclaim. The implementation also distinguishes memory types and uses feedback to balance anonymous and file-backed reclaim. It is still an approximation: access bits are sampled, mappings can be shared, and a page may become hot after its last observation. A generation is not a wall-clock expiration timer, nor does it mean every page was accessed exactly once during a fixed interval.
MGLRU operates within memory-management boundaries such as NUMA nodes and memory cgroups. The same physical host can therefore contain independent reclaim domains. A memory-cgroup limit can force reclaim for one workload while the host still has available memory. Conversely, global pressure may affect several cgroups even if each individual limit appears generous. Compare counters at the same scope as the workload before drawing conclusions.
The feature is conditional on kernel build options and runtime state. Read the running kernel’s own documentation and exposed files rather than assuming every distribution enables the same behavior. The documented controls live below /sys/kernel/mm/lru_gen/ when supported. A missing file can mean the kernel lacks the feature, its configuration does not expose the interface, or the filesystem is not mounted as expected; it is not evidence that a package is missing.
Establish a baseline before changing reclaim
Record the kernel release, boot parameters, memory topology, swap configuration, and cgroup version. Then collect repeated observations during the actual workload, including a quiet baseline and the pressure interval. Useful read-only checks include:
uname -r
cat /proc/meminfo
cat /proc/pressure/memory
cat /proc/vmstat
cat /sys/kernel/mm/lru_gen/enabled
The last path may not exist. Do not turn a failed read into an instruction to create it. MemAvailable is an estimate, not a guarantee that a workload can allocate its next large object. PSI reports time during which tasks were stalled on memory pressure; it does not identify the exact victim page or prove that MGLRU caused the stall. /proc/vmstat fields such as scan, steal, and refault counters are useful as deltas over an interval, not as isolated lifetime totals.
For cgroup-scoped workloads, record memory.current, memory.max, memory.events, memory.stat, and the cgroup’s PSI file when available. Keep the test scope consistent: host-wide reclaim counters cannot directly explain one container’s working-set behavior. Align samples with request latency, throughput, major faults, application cache hit rate, and OOM events. Avoid sampling every millisecond; excessive collection can add noise and overhead.
Runtime control is a policy change
The enabled control exposes a bitmask for supported MGLRU components. The documented default depends on kernel build configuration, and accepted writes do not prove that a component is active on all hardware. Read the value before and after any change, record it with the kernel version, and consult the running kernel’s documentation for valid bits. Do not copy an example value blindly across releases or assume that the hexadecimal value is a universal tuning recommendation.
A controlled test can write the documented enable value only when the host owner has approved the experiment and a rollback plan exists. Such a change alters global reclaim behavior and can affect unrelated workloads. Prefer a canary with representative memory pressure, defined stop conditions, and a known way to restore the prior value. Treat persistent boot-time configuration as a separate change from a temporary runtime test.
MGLRU also documents a min_ttl_ms control intended to protect a recent working set from eviction. This is not a free cache pin. Raising the threshold can preserve more pages and increase the chance that the system reaches OOM when memory cannot satisfy the protected working set. The kernel documentation describes this as a pressure-relief tradeoff, not a substitute for capacity planning. Use it only after measuring the workload and understanding who can be killed if the protection cannot be honored.
The debugfs interface under /sys/kernel/debug/lru_gen is explicitly more experimental than the stable sysfs controls. It supports generation manipulation and proactive reclaim operations for specialized management workflows. Mounting debugfs or writing debug commands on a production host should not be an exploratory first step. Commands can force scanning or evict pages, and syntax or behavior can evolve. Capture the exact kernel release and consult matching documentation before considering these controls.
Compare outcomes, not a single counter
Design an A/B comparison with the same workload, data set, memory limit, swap policy, and observation window. Record request-level latency percentiles, throughput, memory PSI, major faults, application cache misses, reclaim scan/steal deltas, and OOM events. Re-run enough times to distinguish normal variance from a policy effect. A test that reduces resident memory while doubling tail latency is not automatically an improvement.
Keep workload phases separate. Startup, warm-up, steady-state traffic, batch compaction, and idle periods generate different access patterns. The page cache can make a second run look faster even when reclaim policy is unchanged. If possible, use a disposable host or isolated canary and document whether the benchmark begins with warm or cold caches. Do not drop global caches on a live production machine to manufacture a test condition.
When MGLRU is active but pressure remains severe, diagnose the source rather than toggling controls repeatedly. Check for an unexpectedly small cgroup limit, memory growth, swap exhaustion, pinned or unevictable pages, a cache with poor reuse, and NUMA locality problems. A high file-cache figure alone is not a defect; reclaimable file pages can be useful. Likewise, low free memory alone is normal on many Linux systems and does not prove thrashing.
Operational runbook and acceptance criteria
For a reproducible change record, capture:
- kernel release, distribution build, architecture, relevant configuration, and boot parameters;
- whether the MGLRU interface exists and the exact initial and final
enabledvalues; - host and cgroup memory limits, swap availability, and workload identity;
- baseline and candidate intervals with PSI, reclaim deltas, faults, OOM events, and service latency;
- the exact runtime write, authorization, rollback value, and rollback test;
- a conclusion tied to service objectives, not only to memory consumption.
Accept a change only if it improves the target metric or resource objective across repeated representative runs, does not breach latency or availability objectives, and has a verified rollback path. If the kernel lacks the interface, record that fact and stop; changing kernels is a separate compatibility and maintenance decision. If MGLRU is already enabled, do not infer that it is malfunctioning from one burst of reclaim. Correlate pressure, workload access patterns, cgroup boundaries, and application-level effects.
The safe mental model is that MGLRU is a kernel reclaim strategy whose quality must be evaluated under the system’s real constraints. Observe first, change one variable on a canary, measure both memory and service behavior, then either keep the change with evidence or restore the previous state.
Related:
- Linux Pressure Stall Information: Measuring CPU, Memory, and I/O Contention Directly
- Linux KSM: Measure and Operate Kernel Memory Deduplication
Sources: