Skip to content
LinuxHow-To Published Updated 7 min readViews unavailable

Linux MADV_COLLAPSE: On-Demand Transparent Huge Pages

Use MADV_COLLAPSE carefully to request transparent huge-page collapse, understand its synchronous memory costs, verify results, and avoid misleading THP benchmarks.

Linux madvise(..., MADV_COLLAPSE) asks the kernel to collapse eligible memory in a process’s existing mapping into transparent huge pages (THPs). Unlike an ordinary MADV_HUGEPAGE hint that may be acted on later by fault-time allocation or khugepaged, MADV_COLLAPSE is a best-effort synchronous operation over the range’s current state. It can therefore perform allocation, reclaim, compaction, fault-in, and copying work in the caller’s path.

This makes it useful for controlled experiments or applications that have measured a benefit from large-page mappings and can deliberately schedule the cost. It is not a harmless warm-up call, a reservation, or a persistent promise that the mapping will remain huge. It can fail, block, or increase memory use, and any benefit depends on the CPU, workload, page-table layout, kernel, NUMA placement, and virtualization environment.

Know what collapse does and does not guarantee

MADV_COLLAPSE was added in Linux 6.1. The current Linux manual page describes support for private anonymous, shmem, and file-backed pages. The kernel’s transparent-hugepage documentation notes that current madvise_collapse only attempts PMD-sized THPs; do not assume the operation produces every smaller multi-size THP (mTHP) order available on a system. Page size and PMD geometry vary by architecture, so do not hard-code a 2 MiB assumption into portable tooling.

The advice applies to the mapping state at the time of the call. It does not persistently mark the range for future collapse and does not guarantee that later allocations, remaps, writes, or memory pressure preserve a huge mapping. It also does not promise improved application performance: fewer TLB misses may help some memory-intensive workloads, while larger allocation/copy work, fragmentation, memory overcommit, and page-table splits can hurt others.

The kernel attempts eligible hugepage-aligned regions in the requested interval and can continue past regions that fail. It clamps the range to hugepage boundaries. Nonresident pages may be faulted in or swapped in before copying; unmapped pages within a candidate range are zero-initialized in the replacement page, but each candidate hugepage-sized region must have at least one page currently backed by physical memory. A return of zero means all relevant regions were either collapsed successfully or were already PMD-mapped THPs; it does not prove that every other mapping of the same underlying memory is huge.

Distinguish it from THP policy and hugetlb

MADV_HUGEPAGE is advisory and participates in ordinary THP policy. MADV_COLLAPSE requests an immediate best-effort action instead. According to current kernel documentation, collapse is independent of /sys/kernel/mm/transparent_hugepage eligibility and allocation settings and ignores a tmpfs mount’s huge= setting for tmpfs mappings. As a result, setting THP policy to never is not a reliable way to prevent an application from using MADV_COLLAPSE; policy owners should account for the syscall explicitly.

This is also different from hugetlb pages. Hugetlb uses explicitly managed pools and mapping rules, and applications may need reserved huge pages. THP is integrated with ordinary virtual memory and can be split or reclaimed according to kernel policy. Do not interpret successful collapse as a reservation of future huge pages or as a guarantee that a later allocation will succeed.

Call it only on a measured, bounded range

The example below assumes region is the page-aligned start of a valid mapped range. The helper intentionally makes the potentially expensive call explicit at the call site. It does not silently retry or turn failure into a success result.

#define _DEFAULT_SOURCE
#include <errno.h>
#include <stddef.h>
#include <sys/mman.h>

int collapse_measured_range(void *region, size_t length) {
    if (region == NULL || length == 0) {
        errno = EINVAL;
        return -1;
    }

    /* Linux 6.1+; may fault in, reclaim, or compact memory synchronously. */
    return madvise(region, length, MADV_COLLAPSE);
}

Call it outside latency-sensitive request handling unless measurements prove that synchronous work is acceptable. Bound length, validate that the range remains mapped for the call, and coordinate with other threads that could unmap or replace the mapping. The address must satisfy madvise() alignment requirements. Production code should propagate the return value and errno, include a fallback path, and record how often collapse fails; repeated retries under memory pressure can amplify the problem.

Failures have distinct operational meanings. A kernel older than Linux 6.1 or a build without the relevant support may reject the advice. ENOMEM can indicate that a huge page could not be allocated, while a range that includes unmapped addresses may also produce ENOMEM. EBUSY can indicate a memory-cgroup huge-page charge failure. Check the exact error and deployment kernel rather than mapping every failure to “THP disabled.” A successful call is not a performance signal by itself.

Budget reclaim, compaction, and cgroup memory

The kernel may enter direct reclaim or compaction while allocating the replacement huge page, even when ordinary global THP sysfs policy would not have selected that range. That work can stall the calling thread, contend with unrelated allocations, and affect tail latency. A cgroup memory limit can block the huge-page charge even if the host appears to have free memory; inspect the process’s cgroup and relevant memory events along with system-wide memory state.

Some covered pages may be nonresident and have to be faulted in first. On a swapped workload, collapse can turn a seemingly local memory-layout request into storage I/O plus allocation. Avoid calling it over oversized ranges during startup or failover without capacity testing. A server that collapses all worker arenas simultaneously can create a burst of memory demand and compaction pressure.

The replacement may require contiguous physical backing at the chosen huge-page order. Fragmentation and NUMA placement therefore matter. On multi-node systems, the kernel documentation says a collapse allocation is placed on the node that provides the most native pages for the relevant region. This is not a substitute for designing the process’s memory policy and CPU placement. Measure remote memory access and locality after collapse; reduced TLB pressure can be offset by worse NUMA behavior.

Partial mappings and concurrency need attention. Another thread can change protection, unmap, remap, or write within a region while a management operation is being planned. Synchronize application ownership around the range, but do not assume a userspace mutex can constrain unrelated processes sharing a mapping. Re-read mapping and residency evidence after the call, and treat a concurrent mapping change as an expected failure mode in the control path.

Verify at the right level

For a process-level inspection, /proc/PID/smaps reports mapping details including AnonHugePages for anonymous PMD-sized THPs and relevant file-mapping fields such as FilePmdMapped. Reading smaps can be expensive, so use it during a controlled diagnostic window rather than scraping every process at high frequency. The system-wide /proc/meminfo fields AnonHugePages, ShmemPmdMapped, and ShmemHugePages describe different categories; do not treat one as a complete per-process success metric.

The kernel exposes counters in /proc/vmstat for THP fault allocation, collapse allocation, fallback, split, and other transitions. Counters are cumulative: take before-and-after samples and account for other processes on the host. Per-size statistics may also exist under /sys/kernel/mm/transparent_hugepage/hugepages-<size>kB/stats/; feature-probe paths because sizes and files differ by kernel and architecture. khugepaged/pages_collapsed is useful as a rough progress indicator, not an exact count of useful application mappings.

A benchmark should compare equivalent runs with and without collapse. Measure p50/p95/p99 request latency during the call and afterward, startup time, CPU use, memory footprint, page faults, compaction/reclaim activity, TLB behavior where available, and NUMA locality. Use a representative working set that exceeds cache and include the actual memory cgroup limits. Keep unrelated kernel, allocator, and THP-policy changes fixed. A microbenchmark that measures only steady-state pointer traversal misses the one-time collapse pause and memory cost.

Roll out with a safe fallback

Start with one disposable process or bounded mapping and confirm the kernel version, mapping type, alignment, cgroup limits, and fallback behavior. Use a canary that can continue with base pages when collapse fails. Set an explicit time budget around the operation if the application has a latency contract, and do not block all workers waiting on a global collapse phase. Capture kernel, CPU architecture, THP sysfs state, page-size information, cgroup configuration, and before/after counters with the benchmark results.

Do not “roll back” by blindly changing global THP knobs or remapping live memory. MADV_COLLAPSE itself does not create a persistent advice policy; later kernel actions may split or reclaim the resulting THP, and ordinary mapping operations can change page-table shape. If the experiment causes latency or memory regressions, stop issuing new collapse requests, revert the application feature flag, let the process continue under the existing memory policy, and restart only if the application needs a clean baseline. Coordinate any system-wide THP policy changes separately with the platform owner.

Use MADV_COLLAPSE only when a measured workload justifies its synchronous cost and the application can tolerate partial success. Treat success as a memory-layout outcome to verify, not proof of a speedup, and retain a tested base-page path for older kernels, resource limits, and fragmented hosts.

Related:

Sources:

Comments