Linux KSM: Measure and Operate Kernel Memory Deduplication
Use MADV_MERGEABLE and Linux KSM deliberately: scope candidate memory, read the right counters, weigh NUMA and copy-on-write costs, and plan safe rollback.
Kernel Samepage Merging (KSM) is Linux’s opt-in mechanism for deduplicating identical anonymous memory pages. A process marks a mapped range as a candidate with madvise(..., MADV_MERGEABLE); when KSM is configured and its background scanner is running, ksmd searches eligible ranges and can replace identical pages with one shared, write-protected page. A later write gets a private copy through copy-on-write (COW).
This can reduce physical memory use for workloads that hold many stable copies of the same data, including some virtual-machine and application-server workloads. It is not a generic memory compressor or a guarantee that a process’s RSS will shrink. Scanning consumes CPU, per-page bookkeeping has a cost, and writes to shared pages pay COW work. Treat it as a measured workload optimization, not a default switch to turn on everywhere.
Confirm support and understand the scope
KSM requires a kernel built with CONFIG_KSM. The kernel documentation dates the feature to Linux 2.6.32. The application-facing interface is Linux-specific madvise() advice. KSM operates on eligible private anonymous memory, not ordinary file page-cache pages; marking a file-backed range does not turn its cached file data into KSM candidates.
An application can register a range even while the KSM scanner is stopped. Registration is only an eligibility request: it does not itself start ksmd, immediately scan the range, merge any pages, or promise a particular amount of savings. A range that changes frequently or contains mostly unique data can be a poor candidate. Scope advice to the specific stable structures likely to repeat rather than marking every allocation indiscriminately.
madvise() works on whole pages. The address must be page-aligned and the length is rounded up to a page boundary. Keep the mapping alive for the advice lifetime, ensure the requested interval really belongs to the intended mapping, and avoid integer overflow when calculating its end. MADV_MERGEABLE and MADV_UNMERGEABLE are available only when the running kernel has KSM support; unsupported advice normally fails with EINVAL.
Mark only a deliberate region
The following helper expects its caller to pass the base address returned by a page-aligned mapping such as mmap(). It registers the range but does not allocate it, start KSM, or wait for any merge to occur.
#define _DEFAULT_SOURCE
#include <errno.h>
#include <stddef.h>
#include <sys/mman.h>
int mark_candidate_pages(void *page_aligned_region, size_t length) {
if (page_aligned_region == NULL || length == 0) {
errno = EINVAL;
return -1;
}
/* MADV_MERGEABLE is Linux-specific and requires CONFIG_KSM. */
return madvise(page_aligned_region, length, MADV_MERGEABLE);
}
Unlike posix_fadvise(), madvise() reports failure as -1 and sets errno. Production code should preserve and report the error, distinguish unsupported kernels from resource failures, and decide explicitly whether KSM is optional. A failed hint must not silently break application correctness. Never assume madvise() success means memory has been merged: it records advice, while scanning and deduplication happen asynchronously.
KSM may need additional virtual-memory metadata to split or track ranges. The kernel documents possible ENOMEM when the process would exceed vm.max_map_count, and EAGAIN when internal memory is unavailable. A range that contains unmapped gaps can also fail even if some mapped portions were processed. Validate the mapping boundaries and surface errors instead of repeatedly retrying a malformed or oversized range.
Measure value, not just the shared-page count
The KSM sysfs directory, when present, exposes scanner controls and statistics under /sys/kernel/mm/ksm/. Read the availability and values on the actual deployment kernel; distributions may differ in configuration and supported controls. Relevant counters include:
pages_shared: KSM pages currently used as shared backing pages.pages_sharing: additional mappings sharing those pages; this is a useful indicator of potential pages saved.pages_unshared: unique candidate pages repeatedly considered but not merged.pages_volatile: pages changing too quickly to be placed in KSM’s stable tracking structure.pages_scanned,full_scans, andpages_skipped: scanner activity and progress.general_profit: the kernel’s approximate system-wide accounting of memory saved minus KSM metadata cost.
Interpret these as a time series, not a one-time success number. A large pages_shared value alone does not tell you the net benefit: one shared page may have only two mappings, while scanning many unshared pages consumes CPU and metadata. The kernel documentation describes general_profit as an approximate estimate because KSM needs reverse-mapping metadata for candidate pages. Compare it with host memory pressure, CPU use, scan duration, workload latency, and the application’s own memory footprint.
The kernel also exposes KSM-related events in /proc/vmstat. cow_ksm counts copy-on-write events when a process writes to a KSM page; ksm_swpin_copy counts copies associated with swap-in. Rising values are not automatically faults, but if they track an increase in write-heavy workload latency, the mergeable range may be too broad or too mutable. Correlate counters with workload phases and application-level latency histograms rather than attributing every COW to a regression.
Use read-only inspection first. For example, this shell fragment prints the KSM controls and counters that exist on the current host without changing them:
for name in run pages_to_scan sleep_millisecs pages_shared pages_sharing \
pages_unshared pages_volatile pages_scanned full_scans \
pages_skipped general_profit merge_across_nodes advisor_mode; do
file="/sys/kernel/mm/ksm/$name"
if [ -r "$file" ]; then
printf '%-24s %s\n' "$name" "$(cat "$file")"
fi
done
grep -E '^(cow_ksm|ksm_swpin_copy) ' /proc/vmstat || true
Missing files are not a reason to create them: they may indicate that the feature is absent, disabled, or that a particular control is unavailable in this kernel. Do not copy tuning values from another distribution without checking the current kernel documentation and measuring the result.
Operate the scanner deliberately
The /sys/kernel/mm/ksm/run control has three documented values: 0 stops the scanner while keeping pages already merged; 1 runs it; and 2 stops it and unmerges currently merged pages while leaving advised regions registered for a later run. The kernel documentation lists 0 as the default in the usual sysfs-enabled configuration. These are host-wide administrative controls, not per-process switches. Changing them requires appropriate privilege and affects every eligible process on the host.
pages_to_scan and sleep_millisecs influence how much work the scanner performs and how long it sleeps between batches. More aggressive scanning can discover savings sooner but can use more CPU. Start with the distribution’s current defaults, capture a baseline, and adjust only in a controlled environment with an explicit CPU and latency budget. Some newer kernels expose scan-time advisor controls; feature-probe the relevant sysfs files, because the available knobs depend on kernel version and configuration. Do not assume that a value accepted on one host exists on another.
KSM metadata and scan cost can outweigh the saved pages when the candidate population has few duplicates. If pages_unshared is persistently high relative to pages_sharing, or general_profit remains unfavorable, narrow or remove the advised regions. Re-evaluate after application upgrades because allocator layout and data initialization patterns can change duplicate-page yield.
Account for NUMA and write behavior
On NUMA systems, /sys/kernel/mm/ksm/merge_across_nodes controls whether KSM may merge pages from different NUMA nodes. Allowing cross-node sharing can increase deduplication, but the resulting shared page can affect memory locality. Restricting merging to the same node can improve locality on systems with significant NUMA distances, at the cost of some sharing opportunities. This is a workload and topology decision: compare memory savings, remote-memory access, tail latency, and placement under the real scheduler and memory policy.
The kernel requires there to be no KSM shared pages when changing merge_across_nodes; the documented procedure uses run=2 to unmerge before changing it. This is a disruptive host-wide experiment, not a casual online tuning step. Plan capacity for the unmerge and subsequent private copies before attempting it.
Each write to a shared page may need a private page allocation and copy. A workload with large mostly-read immutable tables can behave differently from one that rapidly mutates the same structures. Measure both the steady state and startup or reconfiguration bursts. Include copy-on-write activity, reclaim, swap, and memory-pressure behavior in the test; a memory saving that creates a tail-latency spike during a write burst may not be a useful trade.
Roll out and roll back safely
Use a staged experiment: record baseline memory, CPU, latency, and NUMA behavior; mark one bounded, stable application range; confirm that KSM actually scans and shares it; and compare outcomes over a representative workload cycle. Keep control groups or hosts without KSM advice for a valid comparison. Avoid simultaneous changes to allocator settings, huge pages, swap policy, or NUMA placement, which make attribution difficult.
Removing application advice with MADV_UNMERGEABLE asks the kernel to restore private copies for pages merged in that range. This can suddenly require more memory than is available. The kernel warns that the request can fail with EAGAIN and can also trigger the OOM killer under severe pressure. Do not treat unmerge as a harmless cleanup call. Check available capacity, stop admitting new memory-heavy work if necessary, and perform rollback in a window where the host can tolerate the extra private pages.
Similarly, host-wide run=2 unmerges KSM pages across the machine. Reserve it for a planned maintenance or controlled experiment with sufficient memory headroom. Setting run=0 only pauses the scanner; it does not undo existing sharing. Define these three operations distinctly in runbooks so an operator does not confuse “stop scanning,” “remove this process’s advice,” and “unmerge all host pages.”
KSM is most useful when identical, stable anonymous pages are common and the measured memory savings outweigh scanner CPU, metadata, NUMA effects, and COW cost. Keep the optimization opt-in and narrowly scoped, monitor it as a changing workload characteristic, and make rollback capacity part of the design before enabling it.
Related:
- Linux NUMA Memory Placement: CPU Affinity, Memory Policy, and Locality
- Inside the OOM Killer: How Linux Decides What to Kill When Memory Runs Out
Sources: