Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux HugeTLB Pools: Reservation, NUMA Placement, and cgroup Accounting

Plan Linux HugeTLB pools by page size and NUMA node, distinguish reservations from faults, and validate hugetlb cgroup limits before deployment.

HugeTLB pages are explicitly reserved large pages intended for applications that need a predictable large-page mapping. They are not the same mechanism as Transparent Huge Pages (THP), which the kernel can promote or collapse opportunistically. A system can have substantial free memory and still fail a HugeTLB allocation because the requested page size, NUMA node, reservation, cgroup limit, or pool capacity does not match.

Capacity planning must separate total pages, free pages, reserved pages, surplus pages, and pages charged to a cgroup. It must also include page size: a pool of 2 MiB pages and a pool of 1 GiB pages have different availability and fragmentation behavior. Treat pool changes as a host memory-allocation decision, not a harmless application knob.

Inspect the pool and page sizes

The kernel exposes per-size HugeTLB statistics and controls under sysfs and procfs. Read the exact size directory rather than assuming a default:

grep -H . /sys/kernel/mm/hugepages/hugepages-*/nr_hugepages
grep -H . /sys/kernel/mm/hugepages/hugepages-*/free_hugepages
grep -H . /sys/kernel/mm/hugepages/hugepages-*/resv_hugepages
grep -H . /proc/meminfo

Attributes and supported page sizes depend on the architecture and kernel configuration. A missing directory can simply mean that page size is unavailable. On NUMA machines, inspect per-node HugeTLB directories as well; aggregate free memory can hide that pages exist on the wrong node.

The pool can be sized at boot or adjusted at runtime subject to reclaim, compaction, and availability. A successful write to a pool control is not the same as a guaranteed future allocation for a process: competing consumers, reservations, and memory hotplug can alter availability. Runtime allocation of large contiguous pages may fail even when the requested count is below total free base pages.

Reservation and fault are different moments

An application may reserve HugeTLB capacity before touching the mapping. A reservation accounts for future page use according to the mapping and pool rules, while a page fault instantiates or maps a page to the process. A mapping can fail at reservation time, fault time, or when backing pages cannot be supplied. Operators should distinguish these stages in application logs.

Shared and private mappings, file-backed hugetlbfs, and anonymous HugeTLB mappings can charge and reserve resources differently. A file size may describe potential capacity rather than pages already instantiated. Do not infer resident HugeTLB usage from file length alone. The application should report the requested size, selected page size, reservation result, and whether each fault succeeded.

The hugetlbfs mount exposes HugeTLB-backed files and can impose filesystem-like limits and permissions. The file is not ordinary page-cache storage and does not behave like an arbitrary tmpfs file. A container can see a mounted hugetlbfs path yet still be unable to allocate because the cgroup limit or host pool is exhausted.

NUMA placement and locality

Large pages are physically contiguous within the relevant allocation constraints, so NUMA placement matters. A process bound to one node may fail even if another node has free HugeTLB pages. A memory policy can influence placement but cannot create missing contiguous pages. Check the process CPU and memory policy together with per-node free and reserved counters.

Use the target workload and topology to choose page size. A larger page reduces page-table overhead for some workloads but consumes larger indivisible allocation units and may worsen fragmentation or waste memory for sparse use. Benchmark translation behavior and latency, not just the number of mapped bytes.

If a service starts before its HugeTLB pool or cgroup is ready, it can fail during boot despite succeeding after manual restart. Order the service after the pool configuration and verify the effective cgroup path. Do not make a service silently fall back to ordinary memory if its latency or isolation contract depends on HugeTLB.

Boot-time pool allocation is often more reliable for very large page sizes because memory is less fragmented early in boot. The exact boot parameters and pool behavior depend on the architecture and kernel release. Reserve only the amount justified by an application plan, and leave headroom for the kernel, initramfs, drivers, and ordinary user processes. An oversized reservation can turn a memory optimization into a system-wide capacity failure.

If using persistent HugeTLB files, plan ownership, permissions, mount timing, and cleanup. A shared file can remain allocated after one process exits if another mapping or file reference exists. A service restart that reuses an old file may inherit stale data or an unexpected reservation. Use a documented cleanup policy and validate file size and mapping state before starting the application.

Do not combine THP success metrics with HugeTLB pool accounting. THP collapse can be opportunistic and can fail without preventing an ordinary allocation, while a HugeTLB allocation is explicit and may fail hard when the reserved pool or cgroup budget is exhausted. Monitor each mechanism separately and identify which mapping type the application actually used.

cgroup v2 accounting and delegation

The cgroup v2 hugetlb controller exposes per-page-size usage and limits, including accounting for actual use and, where supported, reservations. The controller can reject new charges even when the host-wide pool has free pages. Inspect the exact files available for each huge page size and the cgroup that contains the process.

find /sys/fs/cgroup/my-service -maxdepth 1 -type f -name 'hugetlb.*' -print
cat /sys/fs/cgroup/my-service/memory.current

The example path is illustrative and usually systemd owns the hierarchy. The memory controller’s current value is not a substitute for HugeTLB-specific counters. Do not edit cgroup files outside a delegated or manager-controlled scope. Configure limits through the service manager where appropriate, then read back the resulting controls.

In a hierarchy, parent and child limits interact. Moving a process can change its charging context, and lowering a limit below current usage may prevent new allocations without immediately reclaiming already allocated HugeTLB pages. A service restart may not free pages if a shared mapping or another process retains the backing object.

Application and failure behavior

Applications should handle allocation, mapping, reservation, and fault failures distinctly. A signal caused by an inaccessible mapping is not a graceful capacity notification. Keep startup checks before the service accepts work, and expose counters for requested pool size, successful mappings, page faults, and fallback decisions.

Do not use HugeTLB only because a benchmark once improved. Compare CPU page-walk cost, memory footprint, fragmentation, locality, and tail latency under production concurrency. A database’s shared buffer pool or a VM’s guest memory model can create a long-lived reservation that starves other workloads.

When pools are exhausted, identify whether the limit is host-wide, per-size, per-node, cgroup-specific, or a reservation. Capture logs, per-node counters, cgroup values, process mappings, and service history before changing the pool. Expanding a pool without calculating node capacity can force reclaim, reduce ordinary memory, or fail during boot.

Safe rollout and acceptance

Test pool allocation on the same kernel, firmware, NUMA topology, and startup order as production. Use a canary cgroup with bounded usage, monitor ordinary memory and latency, and define rollback to the previous pool. Test a process restart, node imbalance, cgroup exhaustion, and host reboot. Keep a separate recovery path if a large boot-time reservation makes the host unbootable.

Acceptance criteria include page size, pool totals and free counts per node, reservation semantics, cgroup path and limit, startup ordering, error behavior, memory headroom for the host, and measured workload benefit. Record the kernel’s supported page sizes and the application version.

HugeTLB is a precise capacity contract only when page size, NUMA placement, reservation, and cgroup charging are measured together. Treat those dimensions as one allocation plan rather than reading a single free-page counter.

Related:

Sources:

Comments