Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux TCP_USER_TIMEOUT: Bound Stalled Established Connections

Set Linux TCP_USER_TIMEOUT deliberately: understand milliseconds, zero-window stalls, keepalive interaction, state scope, and application deadlines.

A TCP connection can remain established while useful progress has stopped. A peer may disappear behind a network partition, a path may black-hole retransmissions, or a receiver may advertise a zero window and stop accepting data. Linux’s TCP_USER_TIMEOUT socket option lets an application bound how long transmitted data remains unacknowledged, or queued data remains blocked by a zero receive window, before TCP closes the connection and reports ETIMEDOUT.

This is a per-socket failure policy, not a general “TCP timeout” and not a replacement for an application deadline. It does not set the retransmission interval, alter when keepalive probes are sent, or bound the initial connection handshake. Applying it without understanding the failure model can turn a recoverable network pause into a needless reconnect storm.

Choose the timeout for the behavior you need

Separate four different clocks before changing a client or service:

  • Connect deadline: how long the application is willing to wait for connection establishment. TCP_USER_TIMEOUT is effective only in synchronized TCP states, so use a nonblocking connect deadline or the networking library’s connect timeout for the handshake.
  • Application or RPC deadline: how long the caller will wait for an operation, including local queueing, request processing, and response handling. Implement this in the application/protocol layer and propagate it where the protocol supports it.
  • TCP user timeout: after an established connection has data that cannot make progress because acknowledgements are missing or the peer advertises a zero window, how long TCP may keep that condition before failing the socket.
  • Keepalive policy: when an otherwise idle connection sends probes to test liveness. SO_KEEPALIVE must be enabled; TCP_KEEPIDLE, TCP_KEEPINTVL, and TCP_KEEPCNT control its idle delay, probe spacing, and probe count where supported.

These controls interact but are not interchangeable. A TCP acknowledgement proves receipt by the remote TCP stack, not that a server application committed a transaction. Conversely, a quiet but healthy connection may carry no outstanding data and therefore have no unacknowledged-data timer to expire. Use protocol heartbeats or application health checks if the requirement is to detect a silent application, and make their failure budget agree with infrastructure load balancers, firewalls, and service deadlines.

What the Linux option means

TCP_USER_TIMEOUT accepts an unsigned integer in milliseconds. A positive value bounds how long transmitted data may remain unacknowledged, or buffered data may remain untransmitted because the peer’s receive window is zero, before Linux forcibly closes the connection and returns ETIMEDOUT to the application. A value of zero selects the system default. The option can be set in any state, but its timeout policy is effective only in synchronized states such as ESTABLISHED, FIN-WAIT-1, and CLOSE-WAIT; it does not provide a connect-handshake deadline.

The option also does not control retransmission timing. TCP continues to perform its normal recovery and retransmission behavior while the user-timeout budget is running. When SO_KEEPALIVE is enabled, TCP_USER_TIMEOUT can determine when the connection is closed after keepalive failure, but it does not change when a keepalive probe is sent. A listener’s setting is inherited by sockets returned by accept() according to the Linux tcp(7) documentation; services should still inspect and test the actual accepted-socket policy rather than assuming their framework preserves it.

The IETF’s RFC 5482 separately specifies an on-wire TCP User Timeout Option that can advertise timeout advice to a peer. Do not treat a local socket setting as an application-level deadline contract or assume the remote application has adopted the same value. Agree on behavior at the protocol layer and verify support at both endpoints when using any wire-level negotiation mechanism.

Set the option on the socket that owns the policy

This small helper sets a 30-second local limit on an already-created TCP socket. The option is Linux-specific; production code should handle an unsupported option and should set it on the socket whose stalled-data lifetime it intends to bound:

#include <errno.h>
#include <netinet/in.h>
#include <netinet/tcp.h>
#include <sys/socket.h>

int set_tcp_user_timeout_ms(int fd, unsigned int timeout_ms)
{
#ifdef TCP_USER_TIMEOUT
    return setsockopt(fd, IPPROTO_TCP, TCP_USER_TIMEOUT,
                      &timeout_ms, sizeof(timeout_ms));
#else
    (void)fd;
    (void)timeout_ms;
    errno = ENOPROTOOPT;
    return -1;
#endif
}

/* Example policy: fail this established stream after 30 seconds
 * without TCP-level forward progress. Choose from service objectives. */
int apply_example_policy(int fd)
{
    if (set_tcp_user_timeout_ms(fd, 30000) == -1) {
        /* Log errno and apply an explicit fallback or fail startup. */
        return -1;
    }
    return 0;
}

This is a socket API example, not a complete connection-management implementation. Validate the option at startup or connection creation, log the chosen value and socket role, and decide whether an unsupported kernel should fail closed, use a documented fallback, or disable the feature. Do not silently ignore setsockopt() errors: otherwise an operator may believe a bound is active when the kernel is using its default.

For accepted sockets, set the option in the accept path if the value is dynamic per client or per service request. If using listener inheritance, test it on the oldest supported kernel and verify with getsockopt() where appropriate. A shared connection pool may need per-destination or per-request budgets above the TCP layer; one fixed socket value cannot express every caller’s deadline.

Derive a value from the service objective

There is no universally correct 30-second or 60-second setting. Start with the maximum time the service can safely retain an established connection that has made no transport progress. Then account for:

  1. Recovery window: how long a transient path outage, wireless roam, route convergence, or peer pause is allowed to last before reconnecting is preferable.
  2. Upper-layer work: whether the request can be replayed safely. A timeout after a write does not prove that the peer failed to process the request; the application may face an ambiguous outcome and need idempotency keys or reconciliation.
  3. Pool behavior: whether one dead connection blocks a worker, how many connections can fail together, and whether retries can overwhelm a recovering dependency.
  4. Infrastructure timers: NAT and firewall idle expiration, load-balancer limits, service-mesh deadlines, and orchestration termination budgets. These may close a flow before the local user timeout.
  5. Fleet behavior: synchronized short timeouts can cause correlated reconnects. Add bounded exponential backoff with jitter and a circuit-breaker or load-shedding policy at the layer that owns retries.

Compare the value with the application’s request deadline rather than blindly making it equal. A request deadline may expire first and cancel the operation; a longer TCP timeout can still matter for connection-pool cleanup, while a shorter one may fail unrelated work sharing that connection. If a multiplexed protocol carries independent requests over one TCP stream, closing the stream affects every in-flight request, not just the slow caller.

Diagnose a timeout before tuning it

When a connection unexpectedly closes, capture the socket’s state and application logs, then correlate packet and infrastructure evidence. Useful host-level snapshots include:

ss -tinp state established
ss -ti dst 198.51.100.40

The destination above is illustrative. ss -ti exposes TCP information such as retransmission and timer details when available, but output varies with iproute2 and kernel versions. Protect process details in production captures because socket diagnostics may reveal command lines or service topology. Correlate the socket tuple with firewall/NAT counters, load-balancer logs, packet captures on both sides where possible, and the application’s send/receive timestamps.

Interpret symptoms by layer:

  • Repeated retransmissions with no ACK: investigate path loss, peer availability, asymmetric routing, firewall state, and packet capture placement. A shorter user timeout only changes when the application gives up.
  • Persistently zero receive window: determine whether the peer application is reading, whether receive-buffer pressure is real, and whether flow control is operating as designed. Closing sooner may limit resource retention but does not repair the slow consumer.
  • Idle flow removed by a middlebox: compare idle timers and keepalive traffic. A user timeout does not itself generate traffic on an idle socket; keepalive or an application heartbeat may be needed to maintain or detect the path.
  • ETIMEDOUT during connect: inspect the connect deadline and SYN retransmissions instead. The synchronized-state limitation means TCP_USER_TIMEOUT is not the right control for that phase.
  • Remote service committed work but the client reports a timeout: handle the result as unknown, not as proof of rollback. Use idempotent operations, transaction identifiers, or a status-query path.

Avoid changing global TCP retry sysctls to solve one application’s connection-lifetime requirement. Global tuning changes behavior for unrelated sockets and can hide differences between interactive clients, long-running replication, and control-plane traffic. Per-socket policy keeps the decision close to the service that understands the consequence of failure.

Test the failure modes deliberately

Use a staging environment or isolated network namespace and test at least:

  • an established connection with a black-holed path while data is outstanding;
  • a peer that advertises a zero receive window and later resumes reading;
  • an idle healthy connection with SO_KEEPALIVE disabled and enabled;
  • keepalive probe failure with and without TCP_USER_TIMEOUT;
  • a separate stalled connect attempt, confirming the application connect deadline still works;
  • partial request delivery followed by timeout, verifying idempotency and ambiguous-result handling;
  • service restart or failover with a large fraction of pooled connections timing out together.

For each case record monotonic timestamps, kernel and iproute2 versions, option values, TCP state, application error, and whether the remote side applied the operation. Do not use wall-clock jumps to measure transport intervals. Acceptance requires the intended timeout to bound the correct established-flow failure mode without truncating healthy recovery, duplicating writes, or creating an uncontrolled retry storm.

Related:

Sources:

Comments