Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux TCP Listen Queues: Diagnose SYN Backlog and Accept Queue Pressure

Diagnose Linux TCP listener overload by separating incomplete SYN requests from established connections awaiting accept, then tune the measured bottleneck.

When clients report intermittent connection failures during a traffic burst, “increase the TCP backlog” is an incomplete diagnosis. Linux has two distinct queues in the TCP listening path: one for incomplete handshakes and another for fully established connections waiting for the application to call accept(). They have different controls, failure modes, and evidence. Raising a single sysctl may do nothing if the application passes a smaller listen() backlog, or if the accept loop is simply not running fast enough.

This guide treats the listener as a system: the client handshake, per-listener queues, kernel limits, application accept and worker capacity, network namespace, and monitoring all matter. The first goal is to identify which queue is under pressure and whether the traffic is legitimate or hostile. Change queue limits only after collecting that evidence.

Separate the SYN request queue from the accept queue

For a TCP server, the application creates and binds a stream socket, calls listen(), then accepts completed connections. Linux’s listen(2) backlog argument is the maximum length for the queue of fully established sockets that have not yet been accepted by the process. Since Linux 2.2, it no longer means the number of incomplete handshakes. The separate incomplete-request limit is controlled by net.ipv4.tcp_max_syn_backlog.

The SYN request queue contains connection attempts for which the handshake has not completed, typically requests in SYN_RECV. The accept queue contains connections whose handshakes completed but which the service has not yet accepted. A remote client can therefore experience a problem at either stage:

  • SYN queue pressure can lead to dropped or delayed handshake requests, retransmissions, or cookie fallback.
  • Accept queue pressure points toward an application that is not accepting completed connections quickly enough, a burst beyond its configured capacity, or an overloaded listener.

These states should not be collapsed into one vague “backlog full” alert. A server can have a healthy accept queue while under a SYN flood, or an empty SYN queue while a stalled worker pool lets established connections accumulate before accept().

Understand the effective limits

The value passed by the application to listen(fd, backlog) is not necessarily the active limit. Linux silently caps it to net.core.somaxconn. The listen(2) manual documents a default somaxconn of 4096 since Linux 5.4, with 128 on earlier kernels; a distribution may override defaults, so read the running network namespace instead of assuming a value.

The incomplete handshake queue is governed separately by net.ipv4.tcp_max_syn_backlog. Upstream kernel documentation describes it as a per-listener limit, and notes that a request socket consumes memory. Increasing it therefore trades memory and state for more room during handshake bursts. The appropriate limit depends on available memory, connection rate, handshake duration, application capacity, and attack exposure.

Read the active controls without changing them:

sysctl net.core.somaxconn \
       net.ipv4.tcp_max_syn_backlog \
       net.ipv4.tcp_syncookies \
       net.ipv4.tcp_abort_on_overflow

The application must also request a useful backlog. If it calls listen() with 128 while somaxconn is 4096, the effective queue cannot become 4096 simply because the host setting is larger. Frameworks, reverse proxies, service managers, and socket-activation units may set or inherit their own value; inspect the running server’s configuration and runtime evidence. The network namespace matters too: a container may have a different listener and sysctl view from the host.

Observe listener queues and kernel counters

Capture listener state during the incident, not only after it has cleared. ss can show TCP listeners and associated processes; request process details with suitable privileges. Sample the same listener repeatedly to distinguish a brief spike from a queue that remains near its limit:

sudo ss --listening --tcp --numeric --processes

For a listener, Recv-Q is the current accept-queue depth and Send-Q is the configured maximum shown for that listener in common Linux ss output. A Recv-Q repeatedly close to Send-Q is evidence that completed connections are waiting for accept(). It does not directly report the separate SYN queue. Always correlate it with application accept rates, worker saturation, and client outcomes rather than tuning from one snapshot.

Kernel TCP counters provide additional context. nstat reads network statistics and can show absolute counters, including zero values, when supported by the installed iproute2 version:

nstat -az | grep -Ei 'ListenOverflows|ListenDrops|Syncookies|TCPReqQFull'

Take before-and-after samples and compare deltas. Names and availability vary with kernel versions. Counters may be cumulative and network-namespace scoped, so check the namespace used by the affected listener. A historical nonzero value alone does not prove a current outage. A growing ListenOverflows or drop-related count during a reproducible traffic interval is more useful when paired with queue depth, application metrics, logs, and client-side handshake evidence.

In a container, inspect the relevant network namespace instead of assuming that host-wide commands describe the service. For example, with a host-visible process ID, an administrator can run a read-only socket inspection in that process’s network namespace:

sudo nsenter --target 12345 --net ss --listening --tcp --numeric --processes

Replace 12345 with a verified process ID. Namespaces, PID visibility, and capabilities affect what can be observed. Do not use a host aggregate counter to assign blame to one tenant without namespace and application evidence.

Treat syncookies as a fallback, not a capacity plan

TCP syncookies allow Linux to respond to SYN queue overflow without storing normal per-request state for every initial handshake. The upstream IP sysctl documentation describes them as a fallback for SYN flooding and explicitly warns against using them to support a server overloaded by legitimate traffic. If cookie-related counters rise during ordinary production load, determine whether the cause is an attack, slow handshakes, an undersized request queue, or simply more legitimate connections than the service can process.

Do not “fix” a growing accept queue by enabling more syncookies. Cookies address a different queue. Likewise, do not treat a cookie counter as conclusive proof of an attack: legitimate bursts can encounter SYN queue pressure too. Compare source distribution, handshake completion, retransmissions, firewall/load-balancer telemetry, and listener-level metrics. If the upstream documentation’s behavior differs from a vendor kernel build, follow the vendor’s kernel documentation and validate the running sysctls.

The related tcp_abort_on_overflow control is not a general queue fix. Kernel documentation warns that resetting connections when a listener is too slow to accept can harm clients, and recommends enabling it only when the daemon cannot be tuned to accept connections faster. Keep it at the distribution’s policy setting unless a tested design calls for otherwise.

Find why the accept loop is falling behind

When the accept queue grows, first investigate the application path. Check whether the accept loop is blocked, paused by overload protection, starved for CPU, constrained by a worker limit, or waiting on synchronous initialization. Inspect file descriptor and process limits, but do not assume that a high connection count is a descriptor leak. Measure request processing, worker queue depth, event-loop lag, CPU throttling, garbage collection, lock contention, and downstream dependencies. A server can accept promptly and then have an application-level work queue fill; conversely, it can fail before the request reaches application code.

For multi-process listeners, verify how workers share or own sockets. SO_REUSEPORT may distribute incoming connections among multiple sockets, but uneven worker health or configuration can leave one listener overloaded while another is idle. Use the application’s own per-worker metrics and socket inspection to establish whether this is happening before changing worker count or enabling reuseport.

For SYN queue pressure, inspect handshake completion and retransmission behavior, upstream load balancers, packet filtering, and connection rate. A queue can fill because the server is receiving malicious traffic, because clients cannot receive SYN-ACKs, or because a legal burst exceeds the host’s service envelope. Packet captures and load-balancer telemetry can distinguish these cases, but capture only the required scope and protect sensitive traffic data.

Tune as a controlled change

Before increasing a limit, establish the failure mode with time-series data and reproduce it in a load test that reflects connection arrival rate, handshake RTT, connection lifetime, worker behavior, and memory headroom. Determine the application’s requested listen() backlog and the effective somaxconn cap. Then estimate how many additional pending connections the proposed limit allows, the kernel memory implications, and how long the application can safely leave a completed connection waiting.

Change one layer at a time in a reviewed configuration-management or sysctl policy. Avoid copying large values from a generic tuning guide. Verify whether the application must recreate its listener to pick up a changed backlog, and confirm the live value through the socket output and repeated queue samples. Keep a rollback path and monitor ListenOverflows, drops, cookie activity, application accept rate, CPU/memory pressure, and client connection success throughout rollout.

The goal is not to make queues enormous. A queue absorbs short bursts; it cannot create worker capacity or make an overloaded service healthy. If a higher limit only converts fast connection failures into long client timeouts, it has shifted the symptom and may worsen tail latency. Scale or repair the accept path, apply load shedding, use an upstream queue where appropriate, or limit abusive traffic at the edge.

Production acceptance criteria

Document which queue was saturated, how it was measured, the listener’s requested and effective backlog, the namespace and kernel version, and the application change or sysctl change made. Validate behavior under a repeatable burst: connections should be accepted or rejected according to an explicit service policy, queue depth should drain within the latency objective, and the relevant counters should remain within the agreed error budget. Include legitimate client traffic, worker restarts, CPU contention, and the expected security posture.

The key diagnostic rule is to identify whether Linux is waiting for a handshake or waiting for the application to accept an already established connection. Tune that measured bottleneck, and preserve the service’s ability to recover when traffic exceeds capacity.

Related:

Sources:

Comments