Skip to content
LinuxHow-To Published Updated 9 min readViews unavailable

Linux tc netem: Reproducible Network Fault-Injection Labs

Build an isolated Linux tc netem lab to test delay, jitter, loss, duplication, and reordering without changing a production interface or host route.

Network failures are difficult to reproduce from an engineer’s laptop: production packet loss is intermittent, WAN latency varies, and the affected path may be outside the team’s control. Linux tc netem provides a queueing discipline that can inject selected impairments into packets leaving a network device. Used inside disposable network namespaces, it makes a useful fault-injection lab for testing application behavior, protocol recovery, retry policies, and observability without changing the host’s production interface.

netem is not a complete simulator of a carrier, Wi-Fi radio, or congested Internet path. Its results are shaped by kernel timer granularity, qdisc placement, the generated traffic, and the rest of the test topology. Treat it as a controlled test instrument: record the exact configuration, validate what the kernel installed, and do not mistake an impairment command for evidence that a production incident has the same cause.

Keep the experiment off production interfaces

tc qdisc add dev DEVICE root ... attaches a root qdisc to the selected device. That changes packet handling for traffic using that device and may conflict with an existing root qdisc or network manager. Never paste a netem example against eth0, a cloud NIC, or an SSH-facing interface until you have reviewed the active qdisc tree, change window, rollback, and effects on all flows.

A pair of network namespaces joined by a virtual Ethernet pair is a safer starting point. Each namespace has its own network stack and routes; the veth pair delivers frames from one endpoint to the other. The following commands create a private IPv4 link using documentation-only addresses:

sudo ip netns add netem-a
sudo ip netns add netem-b
sudo ip link add veth-a netns netem-a type veth peer veth-b netns netem-b

sudo ip -n netem-a link set lo up
sudo ip -n netem-b link set lo up
sudo ip -n netem-a address add 192.0.2.1/30 dev veth-a
sudo ip -n netem-b address add 192.0.2.2/30 dev veth-b
sudo ip -n netem-a link set veth-a up
sudo ip -n netem-b link set veth-b up

Use unique namespace names and first confirm they do not already exist. These commands require Linux, iproute2, and permission to create namespaces and qdiscs. They do not create a bridge to the host or an external network. Verify the isolation and connected routes before adding impairment:

sudo ip -n netem-a address show
sudo ip -n netem-b address show
sudo ip -n netem-a route get 192.0.2.2
sudo ip -n netem-b route get 192.0.2.1
sudo ip netns list

An isolated link makes accidental impact on the host less likely, not impossible if the setup is modified. Do not move physical NICs into test namespaces: deleting a namespace that still has users can keep it alive, and namespace deletion has lifecycle caveats. Keep the test limited to the newly created veth pair and remove the named lab after all test processes have exited.

Add one controlled impairment at a time

To add a base delay with jitter in both directions, attach a netem qdisc to each veth endpoint. Qdiscs operate on egress; applying the impairment to only veth-a affects packets sent from A to B, not replies sent from B to A.

sudo ip netns exec netem-a tc qdisc add dev veth-a root \
  netem delay 100ms 20ms distribution normal seed 42
sudo ip netns exec netem-b tc qdisc add dev veth-b root \
  netem delay 100ms 20ms distribution normal seed 42

sudo ip netns exec netem-a tc -s qdisc show dev veth-a
sudo ip netns exec netem-b tc -s qdisc show dev veth-b

The example configures approximately 100 ms of base one-way delay with a 20 ms normal-distribution variation in each direction. It is not a promise that every observed packet will have that exact latency: scheduler timing, CPU load, traffic shape, and kernel clock resolution affect measurements. With both directions impaired, a simple request/response round trip includes both one-way delays plus endpoint and protocol processing.

Generate a small baseline and compare it with the impaired run:

sudo ip netns exec netem-a ping -n -c 20 192.0.2.2

ICMP echo is only a reachability and timing probe. For application-level behavior, run a real client and server in the two namespaces and record the client’s latency histogram, timeout classification, retries, and server-side outcome. Use a bounded test duration and avoid launching an unbounded retry loop against a lossy path.

Model loss, reordering, and rate limits carefully

Replace the qdisc settings with an explicit delay and independent random packet loss when that is the experiment. Use change only for the root netem qdisc already created on that veth:

sudo ip netns exec netem-a tc qdisc change dev veth-a root \
  netem delay 100ms 20ms distribution normal loss random 1% seed 42
sudo ip netns exec netem-b tc qdisc change dev veth-b root \
  netem delay 100ms 20ms distribution normal loss random 1% seed 42

One percent random loss means the configured probability is applied per packet; it is not equivalent to dropping one percent of requests or one percent of TCP connections. A request may span many packets, and TCP retransmits loss, so application-level failure rates and latency tails can be very different. Correlated or burst-loss models are available, but choose their transition parameters from a written scenario instead of guessing values that merely “look realistic.” A fixed seed guides random loss or corruption generation for repeatable experiments; it does not make end-to-end results identical across kernels, traffic patterns, or machines.

Packet reordering also needs care. netem requires enough delay relative to packet spacing for later packets to overtake earlier ones; otherwise a configured reorder probability may be difficult to observe. Combining delay jitter and explicit reordering can produce more reordering than intended. Begin with a deterministic gap-based test for protocol handling, then add a justified probabilistic model and inspect packet captures to prove that the impairment occurred.

Rate emulation includes packet overhead and cell parameters for representing some link-layer costs. Its achieved rate is limited by the kernel’s timer granularity and the host’s scheduling, so do not use a laptop qdisc as a certification-grade bandwidth or latency instrument. Measure actual throughput, queue statistics, retransmissions, and application outcomes; state test-host CPU and kernel information with the results.

Direction, qdisc placement, and TCP caveats

The root netem commands above affect egress from each veth endpoint. For functional tests of a client and server in namespaces, this often provides a convenient way to impair both directions. It is not automatically representative of every physical path. In particular, the tc-netem(8) documentation warns that mechanisms such as TCP Small Queues can distort TCP performance measurements when netem is placed only on sender egress; for realistic TCP performance testing, the documented guidance is to place the emulator on the receiver’s ingress path.

Linux ingress does not behave like an ordinary egress root qdisc. A common design redirects ingress through an IFB device and applies the shaping qdisc there, but that adds filters, module availability, and another queue to validate. If the research question concerns throughput, congestion control, or realistic WAN bandwidth, document the direction and placement explicitly and validate the topology against the objective. Do not infer TCP performance characteristics from a simple ping test or an egress-only functional lab.

Netem can be combined with other qdiscs, but composition is not automatically safe or intuitive: its documentation notes that interactions with other qdiscs may not always work because netem uses the skb control block for delay state. A root qdisc hierarchy also affects how packets are classified and queued. Record tc -s -d qdisc show dev DEVICE before and after each change, and avoid mixing production shaping policy with experimental impairment unless the hierarchy has been designed and tested deliberately.

Make each run measurable and reversible

Before every run, write down the hypothesis, direction, impairment, rate, seed, test duration, expected application signal, and stop condition. Capture the installed qdisc configuration and counters, and record baseline behavior before impairment. Use tc -s qdisc show afterward to compare packet, drop, and backlog counters with the generated traffic. If counters do not move as expected, the test is not evidence that the intended impairment reached the application.

Change one variable at a time when diagnosing a regression. Delay-only can expose deadline assumptions; loss-only can exercise retransmission and retry behavior; duplication and corruption test different protocol handling. A combined model is appropriate only when the test specifically covers that combination. Include a no-impairment control run so that unrelated host load or test-harness issues are not misattributed to netem.

When the lab is no longer needed, stop its processes and delete only the two namespaces created for this experiment:

sudo ip netns pids netem-a
sudo ip netns pids netem-b
# Stop any listed test processes before removing their namespaces.
sudo ip netns del netem-a
sudo ip netns del netem-b
sudo ip netns list

Deleting a named namespace removes the name, but the namespace may persist while a process still holds it open. Confirm the lab processes have exited and the namespaces are gone. In automation, use a unique run identifier, trap cleanup on failure, and make the cleanup idempotent. Do not use ip netns delete with --all on a shared host.

Acceptance criteria for a useful fault-injection test

  • The experiment runs on an isolated namespace/veth topology or a reviewed disposable host, not an unreviewed production NIC.
  • The tested direction and qdisc placement are explicit; both directions are impaired only when intended.
  • The effective qdisc and its counters prove that packets were queued, delayed, or dropped as designed.
  • A baseline/control run exists, test traffic is bounded, and application metrics are collected in addition to ping.
  • Stochastic runs record the seed and parameters while acknowledging limits to exact reproducibility.
  • TCP throughput conclusions account for receiver-ingress placement and host timer/scheduler limits.
  • Cleanup is scoped to the test namespaces, all processes exit, and no unrelated namespace or root qdisc is removed.

The value of netem is not that it makes a local test identical to the Internet. It lets a team state a network failure model, apply it in a controlled scope, and verify how software behaves under that model. Production-grade results preserve those boundaries and explain the limits of the emulation.

Related:

Sources:

Comments