Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux Network Page Pool: DMA Sync, Recycling, and In-Flight Ownership

Trace Linux page-pool allocation and recycling across NAPI, skbs, XDP, and DMA while preserving references, sync ranges, and teardown safety.

The Linux page-pool API is a recycling allocator used by network drivers for pages or page fragments that back packet buffers and XDP frames. Reusing memory can avoid repeated allocation, mapping, and cache-management costs on a high-rate receive path. The optimization introduces a strict ownership contract: the driver, networking stack, XDP path, DMA engine, and page pool must agree on who owns each page, when device access has ended, which bytes require synchronization, and when an in-flight reference is released.

Page-pool bugs often appear as rare packet corruption, stale data, use-after-free, leaked pages after interface teardown, or a pool that falls back to the general allocator under load. Diagnosis needs to follow the page from allocation through DMA, packet delivery, recycle, and eventual pool destruction.

Pool scope and execution context

A page pool is typically associated with a network queue or NAPI context. The kernel documentation recommends matching the number of pools to NAPI contexts or hardware queues unless hardware constraints prevent it. This keeps fast-path recycling local and avoids unnecessary locking. If a pool is configured with a NAPI instance, that NAPI context must be the sole consumer of the pool’s fast-path state.

Pool parameters include allocation order, pool size, NUMA node, device, NAPI context, DMA direction, queue index, and synchronization range. These fields describe more than performance preferences: the device and DMA direction affect mappings; NUMA placement affects memory locality; and the NAPI association establishes context assumptions. A single shared pool across unrelated queues can erase the locality benefit and complicate concurrency.

Page pool is not a replacement for the DMA API. The device driver still owns the rules for synchronizing memory for CPU access and for matching mappings to device lifetime. Pooling can keep a DMA mapping associated with a page across reuse, but that is safe only if the direction, device, and synchronization contract are correct.

Allocation choices: full pages and fragments

The API supports full-page and fragment allocation paths. Full pages avoid some struct page cache-line traffic when recycled. Fragments can improve memory utilization when packet buffers are small relative to the page, but add reference accounting and shared-page complexity. Multiple fragments may be returned independently, and the last fragment’s release can perform operations for the underlying page.

Choose the allocation shape from the driver’s receive-buffer size and hardware requirements, not merely from a desire to reduce memory. Consider headroom for protocol headers or XDP metadata, alignment required by the device, and maximum frame size including VLAN or tunnel encapsulation. A fragment too small for a jumbo frame can force fallback allocation or truncation; an oversized allocation can waste memory and increase cache/TLB pressure.

When debugging a driver, record the exact allocator function used, requested size, resulting offset and length, page order, and whether the page was split. A code change from full pages to fragments is an ownership-model change and needs concurrency review, not just a benchmark.

Ownership from DMA to the stack

The receive lifecycle begins when the driver allocates or recycles a page, prepares the device-visible buffer, and posts it to a descriptor ring. The NIC writes packet bytes through DMA. After completion, the driver must establish that device access has finished and synchronize the relevant memory for CPU access before parsing it on architectures that require explicit cache maintenance.

The driver then builds an skb or XDP frame around the buffer and transfers ownership to the relevant networking path. If the packet is later freed, the page can be returned to the pool only through a page-pool-aware path. For skbs, drivers can mark the skb for page-pool recycling with the documented helper when appropriate. If the object is copied, cloned, redirected, queued, or retained by another consumer, the page may outlive the RX poll that first received it.

The pool tracks in-flight pages to know when it is safe to release its own object. The driver must return or detach every allocated page correctly. A pool may remain alive after its net device is destroyed while pages are still held in socket receive queues or other consumers. Interface teardown must stop new allocations, detach or destroy the pool according to the API, and allow outstanding references to drain. A dangling skb reference is not fixed by freeing the pool early.

DMA synchronization ranges

If a pool uses PP_FLAG_DMA_SYNC_DEV, the driver supplies the offset and maximum length that should be synchronized for the device. On recycle paths, the pool can use those parameters; on direct release, the caller passes the size that was touched. Incorrect offset or length can leave stale CPU cache lines or unnecessarily synchronize the whole page.

The documentation emphasizes that synchronization parameters apply to the whole page, even when it is split into fragments. Unless the driver author fully understands the fragment accounting and DMA ownership, using offset zero, the full page size, and the conservative full-range sync is the safe baseline described by the documentation. Optimizing to a smaller range should be done only after proving the device’s access window and validating on non-coherent architectures.

The driver remains responsible for synchronizing pages for CPU access. Enabling a page-pool device-sync flag does not eliminate the CPU-side DMA sync obligation. Test on the actual architecture and device, because a coherent x86 development machine may not expose a missing sync that appears on a non-coherent embedded SoC.

Observe recycling and leaks

When supported by kernel configuration, page-pool statistics expose allocation fast/slow paths, cache refills, empty rings, recycle success, full caches, and pages released instead of recycled. The generic netdev netlink family can report page-pool identity, NAPI association, in-flight count, and memory use on supported kernels. Older drivers may expose counters through ethtool or debugfs. Availability depends on kernel version, driver integration, and config options.

Read-only inventory can begin with:

ethtool -i eth0
ethtool -S eth0
ip -s link show dev eth0
ls /sys/class/net/eth0/queues/

These commands do not guarantee page-pool counters exist. ethtool -S names are driver-specific; generic interface drops do not identify pool recycling failures. For netdev page-pool netlink dumps, use a tool that supports the documented netdev family and report the kernel version and tool version.

Compare counters before and after a controlled traffic interval. A high slow-allocation count can mean the pool was empty, traffic burst exceeded its cache, or pages were not returned promptly. A high released-refcount count can be normal for shared pages; interpret it with in-flight memory and workload. A rising in-flight count after traffic stops can indicate buffers retained by sockets, XDP redirects, or a reference leak.

Common failure modes

Stale or corrupted packet data. Check DMA sync direction, device completion ordering, buffer offset, headroom, and whether the right bytes were synchronized for CPU access.

Packets fail only on non-coherent systems. A missing or incorrect DMA synchronization step can be masked on coherent systems. Validate the DMA API contract on the target architecture.

Page-pool memory grows after device removal. Pages may still be held by skbs or XDP consumers. Inspect in-flight references and ensure every terminal path returns or releases its page.

Low recycle rate under load. Pages may be cloned, copied, redirected, or freed from a context that cannot use the direct fast cache. Some releases are expected to go through a ring or general allocator.

Corruption appears only with fragments. Verify fragment reference accounting, last-fragment behavior, DMA sync range, and that no consumer writes beyond the fragment’s bounds.

Fast path works until queue migration or CPU hotplug. Recheck the NAPI context association, queue mapping, affinity, and concurrent access assumptions after topology changes.

Validation and acceptance

For a driver change, test representative packet sizes, fragmented traffic, high PPS, GRO/XDP paths if supported, queue scaling, NAPI budget exhaustion, and interface teardown while traffic is active. Verify packet checksums and payload integrity end-to-end. Test on both coherent and non-coherent architectures when the driver supports both. Measure allocation slow path, recycle counters, in-flight memory, CPU cost, and throughput rather than relying only on line rate.

Keep a trace of queue-to-pool mapping, NAPI identity, DMA direction, sync range, and page lifetime transitions. The test harness should drain socket queues before asserting the pool fully disappears. During teardown, verify the detached pool stops accepting allocations and eventually releases when outstanding references return.

The page-pool optimization is successful only when it improves allocation/reuse cost without weakening the DMA or reference-lifetime contract. Treat every transition as ownership transfer, every sync range as a device-access boundary, and every in-flight page as a reason the pool may outlive the interface that created it.

Related:

Sources:

Comments