Skip to content
LinuxDeep Dive Published Updated 10 min readViews unavailable

Linux AF_XDP: Build a Correct UMEM and XSK Packet Path

Design an AF_XDP receive and transmit path with correct UMEM ownership, XSKMAP queue routing, wakeups, copy-mode detection, and production-grade loss diagnostics.

AF_XDP is a Linux socket family for moving packets between an XDP program and a userspace application through shared descriptor rings and a registered memory area called UMEM. It can support high packet rates and, on compatible drivers and devices, zero-copy data paths. Those properties are conditional: AF_XDP is not automatically zero-copy, does not bypass the need for an XDP program, and can lose packets when queue selection, buffer ownership, or ring service is wrong.

The hard part is not just creating a socket. A production data path must align the XDP redirect target with the socket’s network device and receive queue, keep every UMEM frame owned by exactly one ring or application stage, provision enough buffers for bursts, handle wakeups, and observe drops. This guide focuses on those contracts and on proving which path the host actually selected.

Understand the packet path

An AF_XDP socket (often called an XSK) is created with socket(AF_XDP, ...) and bound to a network device and queue ID. A loaded XDP program must direct ingress traffic to an XSK through an XSKMAP (BPF_MAP_TYPE_XSKMAP) and bpf_redirect_map(). The socket is eligible only for traffic executing on the matching device and queue. A missing map entry, absent XDP program, or socket bound to another queue means packets will not arrive as the application expects; depending on the program’s fallback action, they may continue through the network stack or be dropped.

The socket’s RX and TX rings carry descriptors, not packet bytes. Each descriptor points to an offset and length inside UMEM. UMEM is userspace-allocated memory registered with the AF_XDP socket. Its FILL ring gives available frame addresses to the kernel for receive; RX returns descriptors for received packets; TX publishes packets to send; and COMPLETION returns transmit frames to userspace after the kernel is done with them. Completion is an ownership signal, not proof that a remote peer received the packet.

FILL and COMPLETION rings are associated with a UMEM and its device/queue binding; with shared UMEM, create the required ring pair for each unique device/queue tuple. RX and TX rings belong to individual sockets. The rings are single-producer/single-consumer. If several workers need the same UMEM, design one producer and one consumer for each shared ring or add explicit synchronization and ownership partitioning. Libbpf’s ring helpers provide optimized access patterns, but they do not make concurrent producers safe automatically.

Route to the intended queue

The following minimal eBPF program uses the incoming RX queue as the XSKMAP key. The userspace control plane must populate that key with an XSK bound to the same interface and queue. When the map has no socket for the queue, it leaves the packet on the ordinary networking path.

#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>

#define RX_QUEUE_SLOTS 64

struct {
    __uint(type, BPF_MAP_TYPE_XSKMAP);
    __uint(max_entries, RX_QUEUE_SLOTS);
    __uint(key_size, sizeof(__u32));
    __uint(value_size, sizeof(__u32));
} xsks SEC(".maps");

SEC("xdp")
int redirect_queued_packets(struct xdp_md *ctx) {
    __u32 queue_id = ctx->rx_queue_index;

    return bpf_redirect_map(&xsks, queue_id, XDP_PASS);
}

char LICENSE[] SEC("license") = "GPL";

This is only the forwarding decision, not a complete loader or application. The third argument to bpf_redirect_map() supplies XDP_PASS as the fallback when the requested map entry is unavailable; choose a fallback that matches the service’s packet policy. This fallback-action behavior is part of the modern kernel interface; the upstream libxdp compatibility notes document kernel 5.3 as the baseline for its XSKMAP lookup path, while distributions may backport features. Treat the map behavior and verifier acceptance as runtime capabilities, not a kernel-version guess. Verify that RX_QUEUE_SLOTS covers the queues the control plane may configure. Do not assume an XSKMAP index can redirect traffic across arbitrary devices or receive queues; the kernel checks the XSK’s device and queue compatibility.

Before debugging packet contents, confirm steering. The NIC may distribute traffic across several receive queues using RSS, while one XSK is bound to only one queue. Inspect the interface’s queue and RSS configuration, the loaded XDP attachment mode, and the userspace bind tuple. If only queue 3 is bound, traffic hashed to queues 0, 1, 2, or 4 may never reach that XSK. Either create a coherent per-queue XSK setup and map, or deliberately steer the target flow to the queue that is serviced.

Model UMEM as an ownership state machine

Treat each frame as having one owner at a time. A receive frame moves from application-free storage to FILL, from FILL to kernel/NIC receive capacity, then to RX, then to application processing. If the application wants to transmit the same bytes, it submits the frame on TX and waits until the corresponding address returns through COMPLETION before reuse. If it wants to receive again, it returns the frame address to FILL after processing. Never advertise the same frame simultaneously to receive and transmit, to two FILL rings, or to independent application workers. The kernel documentation warns that this can corrupt packet data because receive DMA and transmit may operate on the same buffer concurrently.

Size the UMEM pool for the maximum outstanding receive buffers, transmit work, application-held packets, and burst reserve, not just the average packets per second. If the FILL ring runs dry, hardware receive capacity can drain and packets can be dropped before the application sees an RX descriptor. If the application stops consuming COMPLETION, TX frames remain unavailable even after the device is done. Track ring occupancy, free-frame count, application processing backlog, and drop counters together.

Chunk size and headroom are part of the packet contract. A frame must fit the maximum packet the application intends to process, including any multi-buffer configuration and metadata/headroom requirements. In aligned-chunk mode, addresses are interpreted on chunk boundaries; do not use arbitrary offsets as though each were an independent frame. Unaligned chunk mode has different address semantics and must be intentionally configured. Keep UMEM address arithmetic range-checked and reject descriptors whose address and length exceed the registered area before parsing bytes.

Separate XDP mode from zero-copy

AF_XDP has generic XDP-SKB and driver XDP-DRV paths, while copy versus zero-copy describes a different property: whether packet data is copied between kernel-owned buffers and UMEM. XDP-SKB is the compatibility path and uses skb processing; a native driver path can improve performance but does not by itself prove zero-copy. Whether zero-copy works depends on driver, device, queue, memory, and kernel support.

The default bind behavior may fall back to copy mode when zero-copy is unavailable. If zero-copy is a hard requirement, request XDP_ZEROCOPY and treat bind failure as a capability mismatch. If copy mode is required for a controlled comparison, request XDP_COPY. After binding, query XDP_OPTIONS and check XDP_OPTIONS_ZEROCOPY; do not infer the active path from a benchmark name, driver name, or successful socket creation. Benchmark copy and zero-copy as separate configurations and retain a tested fallback if the service can operate without zero-copy.

Ring sizes must be powers of two. Configure only the rings the workload needs: an RX-only collector does not need a TX ring. The kernel documentation recommends XDP_USE_NEED_WAKEUP in common setups; when its producer-ring flag indicates that the kernel needs a kick, make the documented syscall rather than spinning or assuming that publishing descriptors alone wakes the device. For TX, sendto() or poll() can notify the kernel when the flag requires it. Avoid syscalls on every batch when the flag is clear, but correctness comes before reducing wakeups.

Handle bursts, packet size, and multi-buffer frames

A single UMEM frame can hold only what its configured chunk and usable headroom permit. Jumbo packets or deliberately segmented packet layouts may span several frames. Multi-buffer AF_XDP requires XDP_USE_SG on the socket and an XDP program section configured for fragments (commonly xdp.frags). Each descriptor still refers to one frame; the XDP_PKT_CONTD option marks that more descriptors belong to the same packet. If either side lacks the required support or flag, multi-buffer packets can be dropped or treated as invalid.

Probe device support instead of assuming that because ordinary XDP works, AF_XDP zero-copy and multi-buffer also work. The kernel documentation describes querying XDP netlink features and the device’s maximum zero-copy segment count. Test maximum supported frame size, fragment count, copy mode, zero-copy mode, and checksum/offload assumptions independently. A batch of RX descriptors can end partway through a multi-buffer packet, so packet assembly state must survive across batches and all fragments must be retained or released consistently.

Backpressure should be explicit. When the application cannot process packets fast enough, it must decide whether to drop early, apply a bounded queue, or steer traffic elsewhere. Unbounded buffering converts packet-rate overload into memory exhaustion and latency collapse. Expose counters for application drops separately from kernel rx_dropped, rx_invalid_descs, and tx_invalid_descs. A high invalid-descriptor count generally indicates a bug in the application/control plane, not ordinary congestion.

Diagnose no traffic and apparent corruption

Start with passive observations before changing NIC state: inspect ip -details link show dev <interface>, ethtool -l <interface>, ethtool -x <interface>, and bpftool net. These show link/XDP attachment details, channel counts, RSS configuration, and attached network programs where supported. Commands and fields vary by driver and tools version. Do not change queue counts or RSS filters on a production interface as an exploratory diagnostic; those operations can disrupt unrelated traffic.

For a no-packet report, check in this order: link and driver support; XDP program attachment and mode; XSKMAP entries; XSK bind device/queue; actual NIC queue selected by RSS or flow steering; initial FILL supply; RX ring consumption; and per-socket drop counters. Confirm the interface and namespace where the program is attached, and ensure the traffic reaches that interface. A successful bind() only establishes the socket’s configuration; it does not prove that packets are routed to it.

For corrupted data, audit frame ownership before blaming the NIC. Check that each UMEM address is in bounds, that the advertised length fits the frame, that a frame is not on FILL and TX simultaneously, and that shared FILL/COMPLETION rings have single-owner access. Validate multi-buffer continuation flags and ensure the parser waits for the last fragment before treating a packet as complete. Use a known traffic generator and compare sequence numbers or payload checksums at each stage.

Roll out and benchmark as a data plane

Build against current kernel UAPI headers and a supported libbpf version, but probe runtime features because a distribution can backport support and a nominal kernel version is not a complete capability contract. Test under the exact NIC model, firmware, driver, queue layout, MTU, CPU/NUMA placement, and traffic profile used in production. Monitor packets per second, drop reasons, CPU cycles, cache behavior, memory footprint, and p99 processing latency. Include bursty traffic and deliberate worker stalls, not only a saturated single-flow best case.

Operate the XDP program and AF_XDP application as one deployable unit. Make map population and socket creation ordering explicit, publish a socket only after it is bound and ready, and define the fallback action during upgrade or process restart. Verify that detaching the program restores the intended kernel networking behavior. Keep a conventional socket or kernel-networking path when availability matters more than peak throughput.

AF_XDP provides a powerful packet handoff mechanism, but its performance comes with explicit memory ownership and device-specific constraints. A correct queue map, disciplined UMEM state machine, observed copy mode, and bounded overload strategy are the foundation; only then do zero-copy and batching results mean anything.

Related:

Sources:

Comments