Skip to content
LinuxDeep Dive Published Updated 5 min readViews unavailable

Linux Page Allocator Fragmentation: Buddy Orders and Compaction Evidence

Diagnose Linux page-allocation failures with buddy orders, migration types, compaction, watermarks, and page-owner evidence instead of free-memory totals.

Linux can report plenty of free memory while a high-order allocation fails. The page allocator obtains physically contiguous ranges in powers of two, and long-running workloads can scatter allocations so that enough free pages exist in aggregate but not in the shape a request needs. NUMA placement, memory zones, watermarks, migration types, pinned pages, and reclaim all affect the result.

Fragmentation is not the same as total memory exhaustion. A request for one page can succeed while a request for a large contiguous block fails. Transparent Huge Pages, some device buffers, and contiguous-memory consumers can expose this difference. Diagnose the requested order and zone before changing global compaction or reclaim settings.

Read allocator snapshots in context

The kernel exposes several read-only views of allocator state:

cat /proc/buddyinfo
cat /proc/pagetypeinfo
cat /proc/zoneinfo
cat /proc/meminfo

The files can be absent, restricted, or differ by kernel configuration. Collect them close to the failure and repeat over time. buddyinfo groups free blocks by order and zone; a large count of order-zero pages does not imply a higher-order range is available. pagetypeinfo adds migration-type detail, while zoneinfo includes watermarks and zone state.

Use the allocation failure stack and requested order where available. An allocation may fail because it requires a specific zone or node, not because the entire system lacks contiguous pages. Device DMA masks, reclaim context, GFP flags, atomicity, and latency restrictions can limit what the allocator is allowed to do.

Do not calculate a universal fragmentation score from one buddyinfo snapshot. Per-CPU page lists can temporarily hold free pages outside the buddy free lists, compaction may be in progress, and counts can change between reads. Treat the proc files as diagnostic samples rather than a transactional snapshot.

Understand buddy orders and migration types

The buddy allocator stores free blocks in orders, where each higher order represents a larger power-of-two page range. When a block is split, smaller buddies can satisfy lower-order requests. Coalescing requires adjacent compatible free blocks, so long-lived allocations and pinned pages can prevent a large block from forming.

Migration types such as movable, reclaimable, and unmovable influence which pageblocks can be compacted or coalesced. A movable workload can still fragment memory if allocations are pinned or placed in unsuitable pageblocks. CMA reserves areas for contiguous allocations but has its own configuration and placement requirements; it is not a generic cure for all high-order failures.

NUMA adds another dimension. A node can have abundant free memory while the policy-selected node or zone is short on the required order. Compare per-node state, process CPU affinity and memory policy, and the device’s locality. Interleaving across nodes can improve throughput for some workloads but may make a local contiguous request harder to satisfy.

Reclaim and compaction are different operations

Reclaim attempts to free eligible pages by dropping clean cache or writing dirty pages. Compaction migrates movable pages to create larger contiguous free areas. Both can consume CPU and add latency. Compaction cannot move every page, and reclaim cannot free data that remains actively referenced or pinned.

Transparent Huge Page allocation may trigger reclaim or compaction depending on policy, but failure to obtain a huge page can fall back to base pages. A driver or subsystem with a hard contiguous-memory requirement may have different failure behavior. Do not interpret a THP fallback as evidence that all high-order allocations are safe.

The vm sysctl documentation includes fragmentation-related controls such as compaction_proactiveness and defrag_mode. Changing them globally can increase background CPU activity, reclaim, or allocation stalls. First measure whether the failing workload benefits from compaction and test the change on a canary with a rollback path.

Find allocation sources and pinning

Page-owner tracking can record allocation stacks for pages when enabled in the kernel. It has memory and CPU overhead and may need boot-time configuration. On an affected lab kernel, it can help identify long-lived allocations and their allocation sites; it is not a free production telemetry feature.

DMA mappings, long-term pins, page tables, huge pages, and device drivers can retain pages that compaction cannot migrate. Check subsystem-specific counters and driver queues before blaming general memory pressure. A process with a small RSS can still hold pinned pages indirectly through shared memory or registered buffers.

Correlate allocator evidence with kernel logs, memory pressure, cgroup limits, NUMA placement, and service latency. If the failing allocation occurs in an atomic context, the kernel cannot wait for the same reclaim or compaction work that a sleeping task can. Increasing free memory may reduce frequency without removing an invalid allocation context.

Reproduce and validate a proposed change

Use a controlled workload that repeats the allocation pattern and records the order, GFP context if available, zone, node, and latency. Compare before and after a kernel or workload change under the same memory size and NUMA policy. Avoid stress tests on production: forcing compaction or memory pressure can stall critical services and trigger OOM behavior.

If adjusting a vm setting, capture the original value, change one control on a canary, and observe allocation success, p99 latency, CPU cost, swap or reclaim activity, and background compaction. Reboot behavior and distribution defaults can differ. Restore the original setting if the target allocation does not improve or if stalls rise.

Acceptance criteria should identify the original failure order and zone, show that the target allocation succeeds under representative fragmentation, and prove that ordinary workloads remain within latency and memory budgets. A higher aggregate MemFree value is not a sufficient success metric.

The page allocator’s constraint is shape and placement, not only quantity. Read buddy orders and migration types at the failure point, identify immovable owners and NUMA restrictions, and use compaction or reservation features only when the requesting subsystem’s contract requires them.

Related:

Sources:

Comments