Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux RPS, RFS, and XPS: Validate Network Queue and CPU Locality

Tune Linux RPS, RFS, and XPS only after mapping NIC queues, IRQ affinity, application CPUs, flow ordering, and measured network bottlenecks.

Linux networking can distribute packet processing across CPU cores through NIC hardware queues and software steering. RSS is commonly performed by the NIC, while RPS moves receive processing in software, RFS tries to align receive processing with the CPU running the application, and XPS selects transmit queues based on CPU or receive-queue locality. These mechanisms are complementary, but enabling all of them indiscriminately can add inter-processor interrupts, contention, cache movement, or packet reordering risk without increasing throughput.

The optimization target is locality and queue balance, not “use every CPU.” Start by mapping interfaces, hardware queues, IRQs, application CPU placement, NUMA nodes, and measured bottlenecks. Then decide whether hardware already provides enough parallelism or software steering addresses a specific mismatch.

Map the current receive and transmit path

Record the interface, driver, link speed, queue counts, RSS configuration, IRQ affinity, and application CPU placement. Read-only inspection commonly starts with:

ethtool -i eth0
ethtool -l eth0
ethtool -x eth0
ls -d /sys/class/net/eth0/queues/rx-* /sys/class/net/eth0/queues/tx-* 2>/dev/null
grep -i eth0 /proc/interrupts

Not all drivers implement every ethtool query, and virtual devices may expose different queue models. Use the actual interface name. Inspect per-queue statistics and drops through the driver’s supported ethtool statistics or netlink interface. A queue count is not proof that traffic is evenly distributed across those queues.

RSS assigns flows to hardware receive queues using a NIC hash and indirection table. Queue interrupt affinity then determines where receive interrupts and initial driver processing run. If the device has enough hardware queues and RSS is balanced, enabling RPS may duplicate work: the stack may enqueue packets to another CPU and send an IPI even though the hardware queue already delivered them to a suitable CPU.

Keep CPU topology in the analysis. Two logical CPUs can share a cache; queues may be attached to a NUMA node different from the application’s memory. A steering map that spreads work broadly can reduce one core’s utilization but increase remote memory access and cache traffic.

RPS: software receive distribution

RPS hashes a packet flow and chooses a CPU from the bitmap configured for a receive queue. The packet is queued to that CPU’s backlog for protocol processing, with an IPI used to wake remote processing as needed. RPS can help when hardware queue count is lower than available CPUs or when a single receive queue is a bottleneck. It does not increase NIC hardware interrupt count, but it adds software queueing and cross-CPU coordination.

The per-queue rps_cpus file contains a CPU bitmap. A zero bitmap means RPS is disabled. Read the current map and compare it with the CPU handling the queue’s interrupt:

for f in /sys/class/net/eth0/queues/rx-*/rps_cpus; do
    printf '%s: ' "$f"
    cat "$f"
done

Do not write a guessed bitmap from a different host topology. CPU numbering can differ across machines and hotplug events. Avoid setting all CPUs automatically: on a multi-queue NIC with RSS already mapping queues to CPUs, RPS may add IPIs with little benefit. For a single-queue device, a locality-aware RPS map can help spread upper-stack work, but must be benchmarked.

RPS flow limit is a separate optional mechanism that can favor smaller flows when a target CPU backlog approaches saturation. It is not a general latency guarantee. If one flow dominates, per-flow ordering constraints mean that distributing packets from the same flow across CPUs is not a safe default strategy.

RFS: follow the receiving application

RFS builds on RPS by using a global flow table and per-queue flow table to steer receive processing toward the CPU where a consuming application thread runs. The desired effect is improved data-cache locality. It must preserve flow ordering when the application moves between CPUs: kernel processing cannot immediately switch CPUs while packets remain outstanding on the old target.

RFS requires both the global rps_sock_flow_entries table and per-queue rps_flow_cnt values. These tables consume memory and must be sized relative to active flow counts, not simply the number of open sockets. The kernel documentation’s suggested sizes are workload examples, not universal production defaults. Begin with the actual number of active flows and validate table allocation and collisions under representative load.

RFS can be counterproductive if application threads are pinned or scheduled far from the NIC, if flows are short-lived, or if the workload is already cache-local. If application CPU placement changes rapidly, steering convergence may lag. Measure cache misses, softirq CPU utilization, queue distribution, and end-to-end latency before and after enabling it.

Accelerated RFS requires support in both the NIC/driver and kernel. It asks hardware to steer a flow to a queue associated with the CPU consuming it. Do not infer support merely from a multi-queue device; confirm driver capabilities and the documented ndo_rx_flow_steer path.

XPS: select transmit queues

XPS selects a transmit hardware queue based on either the CPU sending the packet or the receive queue associated with a flow. Mapping CPUs to separate transmit queues can reduce queue-lock contention and improve cache locality for completions. Mapping receive queues to transmit queues can keep the flow’s transmit side near its receive path.

The xps_cpus and xps_rxqs files configure these maps per transmit queue where supported. A single-transmit-queue device has no queue choice, so XPS cannot provide queue selection benefit. The kernel stores a selected queue for a flow to avoid packet reordering. A flow’s queue can change only under conditions indicating it is safe to do so; do not expect a CPU map edit to remap every active flow immediately.

Inspect current maps and compare them with queue completion IRQ placement. Avoid mapping too many CPUs to one queue if the goal is reducing lock contention. Conversely, one queue per CPU is not automatically optimal if the NIC exposes fewer queues, queue memory is remote, or interrupts are consolidated.

Measure before changing maps

Baseline throughput and latency with the real traffic mix. Capture per-queue packet and drop counts, IRQ counts, softirq CPU time, CPU utilization, context switches, and application-level latency. Use the same peer, packet sizes, flow count, CPU affinity, offloads, and NUMA placement when comparing configurations.

Change one mechanism at a time. First verify hardware RSS and IRQ affinity. If a receive queue is overloaded while other CPUs are idle, test RPS. If kernel processing is balanced but cache locality is poor relative to application placement, test RFS. If transmit queue contention is visible, test XPS. Roll back each experiment before beginning a different one so causal attribution remains possible.

Use a canary or maintenance window for sysfs changes. Record the original values and validate readback. Steering can alter CPU utilization and packet timing for all traffic on the interface. Do not tune a remote management interface without out-of-band access; a poor map can make administration unreliable under load.

Failure patterns

Higher CPU use with no throughput gain. RPS may be adding IPIs and queueing to a path already balanced by RSS. Disable the experimental map and compare per-queue distribution.

One CPU remains saturated. RSS indirection may be skewed, the flow count may be too small for per-flow parallelism, or one flow may dominate. RPS/RFS preserve flow affinity rather than splitting one flow arbitrarily.

Latency worsens after enabling RFS. Application CPU placement, table sizing, scheduling migration, or NUMA locality may be poor. Check CPU and cache metrics rather than only aggregate throughput.

Packets appear out of order. Review queue changes, application CPU migration, NIC offloads, and transport retransmissions. Do not remove flow-affinity safeguards to force a more uniform CPU chart.

Queue counters do not change. The interface may not support the feature, the configuration may be disabled, or traffic may use a different device/namespace path. Confirm the exact netdev and supported attributes.

Acceptance criteria

Document the interface and driver, queue/IRQ topology, CPU and NUMA layout, current RSS indirection, RPS/RFS/XPS maps, and baseline traffic profile. Define the improvement target, such as p99 latency under a fixed throughput or throughput at a fixed CPU budget. Accept a change only when repeated tests improve that target without added drops, reordering, or an unacceptable shift in CPU/IRQ load.

The right configuration follows the packet path and the application’s locality. RSS, RPS, RFS, and XPS solve different stages of that path. Use only the mechanism that addresses a measured imbalance, keep per-flow ordering intact, and save a tested rollback map.

Related:

Sources:

Comments