Skip to content
FreeBSDDeep Dive Published Updated 8 min readViews unavailable

FreeBSD netisr: Diagnosing Deferred Network Packet Dispatch and Queueing

Trace FreeBSD netisr packet dispatch, ordering policies, per-CPU queues, drop counters, and workload evidence before changing network stack tuning.

FreeBSD’s netisr subsystem provides a kernel service for directing packets from network interfaces and other packet sources to protocol handlers. A packet source can request direct dispatch or deferred processing. The distinction affects where work runs, how queues behave, and which ordering guarantees apply. It is a useful layer to understand when a high-rate host shows packet drops, uneven CPU usage, or a mismatch between interface counters and protocol throughput.

netisr is not a promise that every inbound or outbound packet passes through one universal queue. Drivers, protocols, and packet directions use different paths, and direct dispatch may be allowed depending on global and protocol-specific policy. Treat it as one part of the packet-processing pipeline. Confirm that the symptom and counters refer to the same interface, protocol, direction, and workload before attributing a drop to it.

The dispatch boundary

At a high level, a network interface driver receives a frame and represents packet data in mbufs. The driver or another packet source may submit the packet to a protocol through netisr. The service identifies the protocol handler and either runs it directly when allowed or defers it to worker processing. Protocol code then performs protocol-specific work such as IP input and socket delivery. Details vary by device driver, protocol, network-stack configuration, and release.

The kernel developer’s netisr(9) manual describes two relevant submission forms. netisr_queue() always invokes the protocol handler asynchronously in a deferred context. netisr_dispatch() may direct-dispatch if global and protocol policy allow it. The _src variants add an opaque source identifier, which may represent an interface number or socket pointer. These are kernel interfaces, not ordinary administrator commands; administrators observe the subsystem through statistics and runtime configuration rather than calling these functions from a shell.

Deferred queues provide a boundary between packet sources and protocol work, but they add queueing and scheduling behavior. Direct execution may reduce handoff overhead for some paths, yet it changes where protocol work executes. The choice is part of the subsystem’s concurrency model, not a universal performance switch. A driver that already has receive-side parallelism, hardware RSS, or interrupt moderation may interact with netisr placement differently from a single-queue virtual interface.

Ordering policy and flow placement

Protocols register a work-placement policy. The current manual describes three forms:

  • NETISR_POLICY_SOURCE preserves source ordering and does not use mbuf flow IDs for work placement.
  • NETISR_POLICY_FLOW preserves flow ordering based on the mbuf header’s flow ID. If a protocol provides a flow-ID mapping callback for packets that lack a flow ID, netisr can use it; otherwise it falls back to source ordering.
  • NETISR_POLICY_CPU delegates placement to the protocol through its CPU-mapping callback.

These policies protect ordering constraints while allowing work to be distributed. Flow ordering matters because packets in one connection cannot be processed in arbitrary order without consequences. A stable flow-to-CPU mapping can keep related work on one worker while different flows are spread across CPUs. That does not guarantee ideal balance: a single elephant flow can still dominate one CPU, a missing or low-quality flow identifier can limit parallelism, and the handler may have additional shared locks or resources.

Do not infer CPU placement from interface name alone. Check whether the hardware driver exposes multiple receive queues and whether RSS or other hash distribution is enabled. Observe CPU activity during a repeatable workload. A packet queue can be balanced while a different kernel lock, application thread, interrupt source, or storage path remains the bottleneck.

Inspect the counters before tuning

On FreeBSD 15.1, netstat -Q displays netisr(9) statistics, including registered handlers and queue-related evidence. Capture a baseline, run a representative traffic test, and capture a second sample. Preserve output with the timestamp and workload instead of clearing counters on a production host. Counter names and output fields can change across releases, so interpret them against the matching netstat(1) manual.

Pair those counters with interface and protocol evidence:

netstat -Q
netstat -w 1 -I igb0
netstat -i -s
netstat -s -p ip
netstat -m

Replace igb0 with the actual interface. netstat -w samples interface statistics over time; netstat -i -s and protocol statistics provide adjacent context, and netstat -m reports mbuf memory statistics. Not every interface or protocol counter has identical semantics. Read the matching device-driver and netstat(1) documentation before interpreting a counter as a packet drop.

Use deltas over a bounded interval. A nonzero lifetime count may be harmless history; a counter that increases during the reproduction is evidence of a current condition. Check whether drops grow at the interface receive ring, netisr queue, protocol input, socket buffer, or application. Similar user-visible symptoms can occur at each boundary, but the corrective action differs. Also record CPU utilization and per-core imbalance so that queue saturation is not confused with a single busy userland process.

The netisr(9) interface defines a per-CPU queue limit for handlers, while warning that effective queue depth may be as much as twice the configured limit because of implementation details. This is not a target value to copy into a tuning file. A larger queue may absorb a burst but can also retain more queued work and increase latency; it does not create more processing capacity. A smaller queue can expose overload earlier. Change a queue limit only after proving that the relevant queue is the drop point and that the workload needs a different burst budget.

A controlled diagnostic workflow

First identify the packet path. Record interface, driver, firmware, link speed, MTU, protocol, direction, and whether the traffic is local, routed, bridged, tunneled, or inside a VNET jail. Confirm the configuration with read-only tools such as ifconfig, netstat -i, netstat -r, and the relevant device logs. Then run one bounded reproduction with a known endpoint and fixed traffic shape.

While traffic runs, sample netisr, per-interface, protocol, and mbuf statistics at the same cadence. Monitor CPU distribution and identify if one worker or queue saturates. Use a packet capture only when needed to establish packet behavior; a capture itself consumes CPU and memory and can drop packets, so document its filter, interface, snap length, and loss counters. FreeBSD’s BPF capture mechanisms are a separate measurement path and should not be mistaken for kernel protocol delivery.

Build a timeline that associates counter deltas with workload intervals:

Interval Capture
Idle baseline netstat -Q, interface counters, protocol counters, CPU state
Reproduction start Exact command or request, timestamp, endpoints and direction
Sustained load Same counters, application goodput/latency, CPU distribution
Load stop Final deltas, recovery time, remaining queues or errors

Repeat with one deliberate control, such as a different receive-queue count or application concurrency, only where the driver and test plan support it. Avoid combining netisr tunables, NIC offload changes, MTU changes, and traffic shaping in the same experiment. Each can affect packet rate and ordering, making attribution impossible.

Interpreting common symptoms

If interface receive errors rise but netisr queue evidence does not, investigate the link, driver, DMA ring, interrupt moderation, and hardware receive path first. If interface counters look healthy but netisr queue drops increase under load, inspect worker placement, protocol queue behavior, and whether the queue is overloaded by bursts or sustained work. If neither counter changes but applications lose throughput, inspect socket buffers, protocol state, peer behavior, routing, and application backpressure.

An elevated CPU on one core with multiple idle cores is not proof that netisr has the wrong policy. The protocol policy may preserve source or flow ordering, the incoming hash may collapse traffic onto a single flow, or the driver may expose only one receive path. Conversely, many active cores do not prove packet processing is balanced: interrupts, protocol workers, soft interrupts, and application threads may be separate sources of CPU use.

A rising queue-drop counter is not itself a reason to raise the queue limit. It establishes that packets were dropped at that queueing layer, but not whether the underlying issue is a traffic burst, too little service capacity, poor CPU affinity, or an unrelated lock bottleneck. A larger limit may trade drop frequency for queueing delay. Decide which outcome matters for the service and validate both throughput and tail latency.

Version-aware operations and acceptance

Kernel tunables and driver defaults evolve. Capture freebsd-version -kru, the kernel configuration, loaded modules, interface driver, and current netisr(9)/netstat(1) manuals before adopting a procedure from another release. A source build can change code without changing a major release label, so include the kernel revision for custom kernels. Do not set unsupported MIBs from a forum post or copy Linux RSS and softirq instructions into FreeBSD without mapping the equivalent mechanisms.

Before changing configuration, define the success condition: lower queue drops during a fixed load, no unacceptable increase in loaded latency, stable CPU distribution, and no regression in packet ordering or service response. Keep a rollback record and change one variable at a time. If the issue occurs only inside one jail or VNET, test that context separately; the kernel manual documents VIMAGE-specific registration behavior and queue purging when a protocol is disabled in a virtual network stack.

netisr is a concurrency and dispatch layer, not a substitute for packet-path analysis. A disciplined diagnosis correlates its queue and work statistics with hardware counters, CPU execution, protocol behavior, and application measurements. That evidence supports a targeted fix and makes it possible to tell whether an apparent improvement came from dispatch changes or from a different workload.

Related:

Sources:

Comments