Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux NAPI Receive Processing: Poll Budgets, Queue Scaling, and Latency

Trace Linux packet reception through NIC queues and NAPI polling, then tune queue placement, budgets, coalescing, and latency with measurable evidence.

Linux NAPI is a networking-core mechanism that moves packet processing from an interrupt-per-packet pattern toward scheduled polling. A device interrupt signals that work is available; the driver schedules a NAPI instance, processes a bounded receive batch, and completes polling when the queue is drained. Under load, this amortizes interrupt and scheduling overhead. Under light load, interrupt moderation and batching still affect latency, so the fastest throughput configuration is not automatically the best interactive or RPC configuration.

NAPI is not synonymous with a particular CPU, hardware queue, or execution context. Multi-queue devices, RSS, interrupt affinity, software steering, threaded NAPI, and driver-specific behavior all influence where packets are processed. Diagnose the whole path before changing a budget or pinning a thread.

Follow a packet from queue to protocol stack

A NIC commonly places received frames into a descriptor ring associated with a receive queue. The driver programs the device and arranges for a hardware interrupt, often MSI-X for multi-queue devices, to announce work. The interrupt handler acknowledges or masks the relevant source and schedules the queue’s NAPI instance. The poll method then reclaims descriptors and builds packet buffers, handing processed packets toward the networking stack, often through Generic Receive Offload (GRO).

When the queue still has work after the poll budget is exhausted, the driver reports the budget as consumed and polling continues. When the queue is drained, the driver completes the NAPI instance and may re-enable its interrupt. Correct ordering matters: re-enabling an interrupt before successful completion can race with new arrivals and strand work. The exact device interrupt and descriptor operations are driver-specific, but NAPI’s scheduling and completion contract is shared.

The conventional path runs NAPI work from network softirq processing. Linux also supports threaded NAPI, where a dedicated kernel thread performs polling, and busy-poll modes where userspace participates in triggering packet processing. These modes change CPU accounting, latency, affinity, and power behavior. Do not assume that seeing a NAPI instance means it always runs in NET_RX softirq context.

Understand the poll budget precisely

The poll callback receives a budget that limits receive packets processed in one invocation. Transmit completion cleanup is not limited in the same way and may run when the receive budget is zero. That zero-budget case exists when the core asks a driver to perform transmit-only cleanup; receive-specific APIs such as page-pool or XDP processing cannot be used in that call, and the driver must not call napi_complete_done() with a zero budget.

When receive work remains, the poll method should return exactly budget so the core polls again. When all events are drained, it calls napi_complete_done() before returning. If the queue completes after processing exactly budget packets, the interface has no distinct return value for “exactly budget, but drained”; the driver must handle that edge deliberately. The following is the documented completion shape, with device-specific cleanup omitted:

if (work_done < budget &&
    napi_complete_done(&queue->napi, work_done)) {
	device_unmask_rx_irq(queue);
	return min(work_done, budget - 1);
}

/* More work may remain, or completion lost a race: poll again. */
return budget;

This is a kernel-driver fragment, not a compilable standalone program. The budget guard prevents calling completion in the zero-budget case. The IRQ is unmasked only after successful completion. A real poll routine must also account for every descriptor, allocation failure, transmit completion, XDP/page-pool rules, and the driver’s interrupt race protocol. Do not copy the fragment without reading the current NAPI documentation for the target kernel and driver model.

The per-poll budget is distinct from global receive softirq limits. Raising one setting may simply move the bottleneck, increase time spent on one CPU, or increase tail latency for other work. Inspect kernel and distribution configuration before tuning any budget.

Separate hardware distribution from software steering

Receive Side Scaling (RSS) hashes packet flow fields in the NIC and selects a receive queue. Hardware queue/IRQ mapping and CPU affinity then influence which CPU handles the interrupt and NAPI work. Receive Packet Steering (RPS) can steer protocol processing in software after the driver receives a packet; Receive Flow Steering (RFS) tries to place processing near the consuming application; XPS steers transmit queues. These mechanisms operate at different points and can interact.

The queue count, RSS indirection table, IRQ mapping, NUMA locality, and application placement should be inspected together. One receive queue per CPU is a common starting shape only when supported and appropriate; a NIC may expose fewer queues, a driver may map NAPI instances differently, and a workload may have a few hot flows that cannot be parallelized by adding queues. More queues can also consume more interrupt and memory resources.

Useful read-only checks include:

IFACE=eth0
ethtool -l "$IFACE"          # channel counts, if supported
ethtool -x "$IFACE"          # RSS indirection and hash configuration, if supported
ethtool -c "$IFACE"          # interrupt coalescing, if supported
ethtool -S "$IFACE"          # driver-specific counters
ip -s link show dev "$IFACE"
cat /proc/interrupts

Driver support and counter names vary. A command that returns “operation not supported” is not proof that the kernel networking stack lacks a capability; it may be absent from that device or driver. Check the NIC driver’s documentation before changing channel counts, affinity, or coalescing. Capture a baseline and preserve a rollback procedure.

Treat batching and coalescing as latency tradeoffs

Interrupt coalescing delays or batches notifications according to device and driver policy. GRO combines packets for upper-stack processing and can reduce per-packet cost. NAPI budgets bound a poll cycle. These mechanisms reduce overhead at high packet rates, but batching introduces a wait or larger burst at lower rates. An aggressive latency profile may improve tail response while consuming more CPU and reducing throughput efficiency.

Measure the service objective that matters: throughput, packet drops, p50/p99/p999 latency, CPU per packet, softirq share, and queue imbalance. Use a controlled workload with packet size, flow count, connection behavior, offloads, concurrency, and CPU placement recorded. A single bulk-flow benchmark can conceal RSS imbalance; many small flows can stress a different path. Include idle-to-burst transitions because a configuration that performs well at saturation may be poor after traffic goes quiet.

Avoid applying a global ethtool -C value copied from another NIC. Hardware may implement adaptive coalescing, different timer units, or no requested mode. Change one variable at a time, observe counters and tail latency, and revert if drops, queue buildup, or latency regressions appear.

Diagnose where packets are delayed or lost

Start with link counters and driver statistics, then compare them to application receive counts. Inspect per-queue packet and drop counters, interrupt distribution in /proc/interrupts, CPU softirq time, and socket-level overflow indicators where available. The kernel’s /proc/net/softnet_stat can help identify networking backlog conditions, but field layout and interpretation should be matched to the running kernel rather than copied from an old dashboard.

If one CPU is saturated while others are idle, inspect IRQ affinity, queue mapping, flow hashing, RPS masks, and application affinity. If queues are balanced but p99 latency rises, compare coalescing, GRO, budget exhaustion, scheduler contention, and application read cadence. If device RX counters rise but socket delivery does not, trace the path between driver receive, protocol processing, routing/firewall, and socket queues rather than assuming NAPI itself dropped packets.

Threaded NAPI can isolate polling in per-instance kernel threads and make scheduling or affinity more explicit, but it does not create extra NIC queues or guarantee lower latency. The current kernel documentation describes enabling it through the netdev sysfs interface or Netlink for a specific NAPI instance. Check support, queue-to-instance mapping, and the target kernel before changing it. Busy polling similarly trades CPU cycles for reduced wait on selected low-latency paths; it is not a general-purpose server optimization.

Define an acceptance test before tuning

For each candidate configuration, record the kernel release and config, driver and firmware, NIC model, queue count, RSS table, IRQ affinities, coalescing, offloads, CPU/NUMA topology, workload, and application placement. Run repeatable tests at idle, ramp-up, steady load, and overload. Track throughput and tail latency beside per-queue drops, interrupt counts, softirq CPU, and application errors.

Accept a change only when it improves the intended service metric without hiding packet loss or merely moving work to a different CPU. Repeat tests after reboot if settings are persisted through distribution networking tools, because a transient ethtool command is not durable configuration. Preserve a known-good baseline and verify the rollback path.

NAPI is a contract between the driver and the networking core, not a single tuning knob. Understanding poll completion, queue placement, batching, and context gives operators a way to explain the measurements before making a risky change.

Related:

Sources:

Comments