Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux Managed IRQ Affinity: MSI-X Queues, CPU Hotplug, and Effective Placement

Map Linux MSI-X queue interrupts to CPUs safely by distinguishing requested from effective affinity, managed IRQ rules, hotplug, and NUMA locality.

Interrupt affinity determines which CPUs may service a device interrupt. On a high-throughput NIC or NVMe device, this decision interacts with queue count, driver design, CPU topology, NUMA memory placement, and CPU hotplug. Reading /proc/interrupts is a useful start, but it does not reveal the full relationship between a hardware queue, an MSI-X vector, the IRQ core’s requested mask, and the CPU that actually receives the interrupt.

The important operational distinction is between ordinary IRQ affinity and affinity-managed interrupts. A user may be able to adjust some IRQs through procfs, while managed IRQ placement belongs to the kernel and driver. An affinity value that appears writable or a queue-to-CPU diagram copied from another server does not prove the device is honoring the desired placement on this host.

From hardware vector to Linux IRQ

Traditional PCI INTx interrupts use shared line-style signaling. MSI and MSI-X instead let a device request message-signaled interrupts; MSI-X provides a table of independently configurable vectors and is commonly used for per-queue interrupts. The actual number of vectors a driver obtains depends on device capability, kernel APIs, available vector space, and driver policy. One queue does not universally equal one vector, and management vectors may be separate from data queues.

The driver requests vectors through PCI APIs and can ask the IRQ subsystem to spread affinity across CPUs. The IRQ core computes affinity masks and associates them with Linux IRQ numbers. A driver using managed affinity also owns a lifecycle obligation: when all CPUs in an interrupt’s allowed mask go offline, the interrupt may be shut down, so the associated queue must be quiesced before that happens. CPU hotplug can therefore expose a driver bug or an unexpected service interruption even when steady-state placement looked correct.

Begin with the runtime mapping rather than inferring it from a PCI function name. /proc/interrupts includes the Linux IRQ number, per-CPU interrupt counts, and device or queue labels supplied by drivers. Device sysfs can expose MSI vectors and queue attributes, but the filenames vary by bus, driver, and kernel. Correlate these sources with the driver’s documentation and workload traffic.

Requested, effective, and managed masks

For ordinary IRQs, /proc/irq/IRQ/smp_affinity and smp_affinity_list describe the permitted CPU mask. effective_affinity or effective_affinity_list, when present, reports where the interrupt can actually be delivered after controller and kernel constraints. A requested mask can therefore differ from effective placement. Some architectures or interrupt controllers do not support arbitrary affinity; a write may be rejected or the effective mask may remain unchanged.

Affinity-managed IRQs are different. Their placement is established through kernel/driver APIs such as pci_alloc_irq_vectors_affinity() and updated through CPU hotplug handling. The user cannot generally override managed affinity by writing procfs masks. The kernel’s documentation also describes isolcpus=managed_irq as a best-effort avoidance hint for managed interrupts, not a guarantee that isolated CPUs will never receive them.

Use a read-only inventory before changing anything:

cat /proc/interrupts
for irq in /proc/irq/[0-9]*; do
    [ -d "$irq" ] || continue
    printf '%s ' "${irq##*/}"
    for name in smp_affinity_list effective_affinity_list; do
        if [ -r "$irq/$name" ]; then
            printf '%s=' "$name"
            tr '\n' ' ' < "$irq/$name"
        fi
    done
    printf '\n'
done

An absent effective_affinity_list file is not proof of a broken kernel; interface availability depends on configuration and architecture. Preserve the output with kernel release, device driver, CPU online mask, and PCI topology. Repeat after a CPU hotplug event because managed masks can change or an interrupt can become disabled.

Queue locality and the rest of the data path

For a multiqueue NIC, interrupt affinity is one stage in a longer receive path. RSS hardware selects a receive queue, that queue generates or schedules work, and NAPI polls packets. RPS/RFS and application thread placement can move later processing to other CPUs. For transmit, XPS and queue selection also affect where work is enqueued. An IRQ pinned to CPU 2 does not guarantee that the application consuming the packet runs on CPU 2.

NUMA locality matters when the device is attached to one node and its queues, DMA buffers, and worker threads are spread across distant nodes. Some drivers allocate queue memory based on IRQ affinity or device locality; others have different policies. Check driver behavior rather than assuming affinity reconfigures all queue memory. Also account for CPU siblings, cache sharing, and other IRQs competing on the selected core.

Capture interrupt counter deltas under representative traffic, then compare softirq time, NAPI poll statistics, queue drops, per-CPU utilization, and application latency. The aggregate count can hide one hot queue or an imbalance caused by a flow hash. If a single flow dominates, evenly distributing IRQs cannot divide that flow across all queues without changing the flow-steering model.

CPU hotplug and isolation boundaries

When a CPU goes offline, ordinary affinity masks are adjusted or the IRQ is moved according to the supported controller behavior. Managed IRQs use a kernel-owned policy; if the last permitted CPU disappears, the interrupt can be shut down until an eligible CPU returns. A driver must stop the hardware queue before the vector becomes unavailable and restart it safely when affinity is restored. Test suspend/resume and CPU online/offline operations on representative hardware, not only a VM with a synthetic interrupt controller.

CPU isolation is also not equivalent to a hard promise that no kernel work will ever run there. The managed_irq isolcpus option only asks affinity-managed interrupts to avoid a mask when another eligible placement exists. If the interrupt’s only possible CPUs are isolated, the hint has no effect. Ordinary IRQ affinity, per-CPU timers, workqueues, RCU callbacks, and application threads require separate analysis.

Do not manually write affinity masks while irqbalance, a device manager, or a driver may be changing them. First identify the owner and current persistence mechanism. If experimentation is warranted, use a non-production interface or test host, record the original masks, change one queue/vector at a time, and verify the effective list plus traffic counters. Restore original settings before reboot or maintenance ends unless a reviewed persistent configuration is intended.

A repeatable measurement plan

  1. Record /proc/interrupts, online CPUs, device NUMA node, queue count, driver version, and qdisc/NAPI settings.
  2. Generate a fixed traffic pattern with both many flows and one dominant flow. Use the same packet sizes and offered load for each run.
  3. Measure per-CPU IRQ deltas, softirq utilization, receive/transmit drops, throughput, and p99 application latency.
  4. Change one non-managed affinity setting in a disposable test and repeat the workload; capture both requested and effective masks.
  5. Offline and online an eligible CPU only in a maintenance test, then verify queue recovery and interrupt delivery.

Do not evaluate affinity by maximizing interrupt counts on one core or minimizing them on all others. A queue can process packets in NAPI context without one hardware interrupt per packet, and interrupt coalescing changes count rates. The desired result is a stable distribution that meets the workload’s latency, throughput, and CPU-isolation goals without queue starvation.

Production acceptance criteria

Document which IRQs are managed, which are user-affined, which service queues they map to, and which tool owns configuration. State the supported CPU hotplug and suspend behavior. For each change, show requested and effective masks, counter deltas, per-queue drops, and application-level latency under the same workload. Include a rollback path and verify that reboot or irqbalance does not silently overwrite the intended state.

If a device fails to recover after CPU hotplug, inspect its managed affinity masks, driver logs, queue state, and kernel IRQ diagnostics before changing the boot command line. An apparently “unbalanced” vector layout may be required by the driver or controller. Avoid pinning every interrupt to one CPU just because that makes one chart look cleaner; the real acceptance criterion is end-to-end processing locality with safe lifecycle behavior.

Related:

Sources:

Comments