Linux Conntrack Capacity: Diagnose State-Table Pressure Safely
Measure Linux Netfilter conntrack occupancy, insertion failures, event loss, and timeout pressure before changing firewall state or table limits.
Linux connection tracking (conntrack) maintains flow state used by stateful firewall rules, NAT, and other networking components. When the table approaches its configured limit, symptoms can look unrelated: new flows fail, NAT mappings are missing, firewall rules classify packets differently, or a node logs conntrack insertion failures. The right response is not to flush the table or blindly raise a sysctl. First confirm the affected network namespace, measure occupancy and insertion behavior over time, identify which traffic creates entries, and then size or tune the state lifetime against a documented load model.
Conntrack state is distinct from a TCP socket’s state. It is a kernel Netfilter record for a packet flow and, where applicable, its reply direction. The entry may outlive an application socket or may expire before the application considers its work finished. A rule such as ct state established,related accept consumes that state but does not itself show that the table has capacity.
Establish the namespace and current owner
Collect evidence in the network namespace where the firewall or NAT path runs. A host command can report a different table from a container, router namespace, or service namespace. On the host, begin with:
sudo ip netns list
sudo ip netns identify "$PID"
sudo readlink "/proc/$PID/ns/net"
sudo cat /proc/sys/net/netfilter/nf_conntrack_count
sudo cat /proc/sys/net/netfilter/nf_conntrack_max
Replace PID with the affected process and compare namespace identities with the firewall process. If the workload uses a named namespace, repeat the count, limit, and flow inspection there:
sudo ip netns exec "$NS" cat /proc/sys/net/netfilter/nf_conntrack_count
sudo ip netns exec "$NS" cat /proc/sys/net/netfilter/nf_conntrack_max
sudo ip netns exec "$NS" conntrack -C
sudo ip netns exec "$NS" conntrack -S
The kernel exposes nf_conntrack_count as the number of currently allocated flow entries and nf_conntrack_max as the maximum allowed entry count. The conntrack utility can report a count and in-kernel statistics through ctnetlink. Read-only inspection is safer than dumping or mutating state, but output can contain source/destination addresses and ports; restrict access, retain it only as long as needed, and redact customer identifiers before sharing.
Check whether the relevant kernel module/features and userspace tools are present before treating a missing path or command as a zero count:
test -r /proc/sys/net/netfilter/nf_conntrack_count && echo "count readable"
test -r /proc/sys/net/netfilter/nf_conntrack_max && echo "limit readable"
command -v conntrack || true
If these files are absent in the target namespace, determine whether conntrack is unavailable, the module is not loaded, the namespace does not use that path, or permissions prevent inspection. Do not load modules or alter firewall configuration during a read-only incident collection without the host’s change process.
Measure a trend, not a single percentage
Sample count and maximum at a steady interval during both normal and peak traffic. The ratio count / max gives a useful occupancy indicator, but one snapshot cannot distinguish a stable high-water mark from a rapidly growing leak or a short burst. Record at least:
- current and peak count, configured maximum, sampling interval, and network namespace;
- new connection rate, close/expiry rate, and protocol distribution where available;
- insert failures, drops, early drops, and kernel log messages over matching timestamps;
- NAT, firewall, container, and load-balancer changes that alter the number or lifetime of tracked flows;
- memory pressure, CPU cost, and hash-table parameters on the same host and kernel build.
For event-level inspection, start with narrowly filtered read-only commands:
sudo conntrack -L -p tcp --state ESTABLISHED
sudo conntrack -L -p udp
sudo conntrack -E -o timestamp --event-mask NEW,DESTROY
The first two can produce very large and sensitive dumps on a busy node; use them only on the target namespace and scope the filters further where supported. An event listener is useful for observing creation and destruction rates, but it is not a durable accounting system unless its loss behavior, consumer lag, and persistence have been engineered. If the utility reports ENOBUFS, its netlink receive buffer or consumer throughput may be insufficient; a larger buffer costs memory and does not fix a permanently slow consumer.
For targeted inspection, filter on a known address, protocol, port, mark, or zone rather than exporting the full table. Compare original and reply tuples with NAT and the firewall’s routing path. A large number of short-lived SYN_SENT, UDP, or other entries may indicate a burst of legitimate fan-out, a retry storm, an address/port churn pattern, or an unexpected packet path. Counts alone do not identify the cause.
Interpret limits and hash buckets correctly
nf_conntrack_max is the table’s maximum allowed number of entries. Kernel documentation states that this limit defaults to nf_conntrack_buckets; it also notes that conntrack entries are represented for original and reply directions, so a full table with defaults has an average hash-chain length of about two, not one. nf_conntrack_buckets controls hash-table size and has a different mutability scope: the kernel documentation specifies that it is writable only in the initial network namespace. Do not assume every conntrack parameter can be changed from a container or that raising the entry maximum alone is free.
More entries can increase memory use and hash-chain work. A mismatch between max and buckets may be intentional, but must be supported by measurement on the actual kernel and traffic profile. Before proposing capacity changes, estimate concurrent flow cardinality using connection arrival rate and observed state lifetimes, including retransmit, idle, UDP, and application-specific behavior. Validate the estimate against peak and failure-mode tests rather than using a generic “connections per GB” rule.
Timeouts are operational state-retention policy. Shortening them can free entries sooner but may break legitimate idle TCP sessions, UDP exchanges, failover, or NAT mappings. Lengthening them may improve tolerance for quiet flows while keeping stale state allocated longer. Defaults and effective values can differ by kernel and distribution; read the target system’s values and compare them with service and middlebox timers before changing anything.
Distinguish capacity exhaustion from other failures
When an application reports connection failure, correlate the time with the kernel log, count/max trend, ctnetlink statistics, and packets on the relevant interfaces. A table-full log paired with rising count near the configured maximum and increasing insertion failures is stronger evidence of capacity pressure than a high ratio alone. A low count with no insertion failures suggests looking elsewhere: routing, firewall verdicts, NAT configuration, namespace mismatch, asymmetric routing, or a missing reply path.
Conntrack statistics can include counters such as inserts, insert failures, drops, early drops, errors, and chain-too-long conditions. Their exact presentation depends on kernel version and tool; compare deltas over time rather than treating a historical nonzero counter as a current incident. A hardware or software flowtable may offload established flows and change what a simple snapshot shows; account for flowtable counters and offload status when the datapath uses them.
Do not assume every packet traverses conntrack. Stateless rules, notrack, bridge paths, hardware offload, and namespace boundaries may change which packets are tracked. The nftables ruleset can show intent, while counters and packet captures show what happened. Review both directions of the flow, the zone if zones are configured, and the order of defragmentation, tracking, NAT, and filter hooks relevant to the deployed rules.
Choose a corrective action from evidence
Use the least disruptive remedy that addresses the measured constraint:
- Unexpected flow creation: find the application, retry behavior, scanner, container churn, or routing loop that creates entries and correct that source.
- Legitimate concurrency exceeds capacity: model the peak and failure surge, validate memory and CPU headroom, then plan a staged, persistent adjustment to the relevant limit and, if justified, bucket configuration.
- Entries persist longer than intended: verify the protocol state and application/middlebox lifecycle before tuning the specific timeout. Test both established long-lived flows and cleanup behavior.
- State is missing between firewall peers: inspect the HA synchronization design and failover timing; do not confuse conntrack-table capacity with replication health.
- The table is full during an incident: preserve logs and counters, prioritize critical recovery paths through the established operational procedure, and avoid deleting arbitrary entries that may carry active NAT or firewall state.
Never use conntrack -F as a generic cleanup. It flushes the selected connection table, invalidates active tracking/NAT state, and can interrupt unrelated flows. Likewise, do not zero statistics casually: it destroys comparison evidence for every observer sharing those counters. If an emergency state deletion is unavoidable, scope it to an approved tuple/zone and document the expected impact and rollback limitations.
Validate the change under normal and surge conditions
Use a staging system or representative test namespace to test expected peak concurrency, short-flow bursts, long-lived sessions, UDP-heavy traffic, IPv4 and IPv6, NAT, and failover. Verify the effective sysctl values after the actual persistence mechanism and after reboot or network namespace recreation. Monitor count/max, insertion failures, latency, packet loss, memory, CPU, and application outcomes together.
Acceptance should show that the peak and recovery surge remain below the planned operating threshold with safety margin, entries expire according to the intended policy, no unexpected drops or insertion failures occur, and firewall/NAT behavior survives restart or failover. “We increased the number and it stopped logging” is not enough if it hides an unbounded flow-creation bug or shifts resource exhaustion elsewhere.
Related:
- How to Configure a Firewall with nftables
- Linux Policy Routing with ip rule: Source-Based Paths and Verification
Sources: