Linux SO_REUSEPORT: Multi-Listener Scaling and BPF Selection
Scale Linux TCP or UDP listeners with SO_REUSEPORT while validating bind order, effective-UID rules, flow distribution, and optional BPF selectors.
Linux SO_REUSEPORT allows multiple sockets to bind an identical local socket address so the kernel can distribute incoming TCP connections or UDP packets among a group of listeners. It is used by multi-process and multi-threaded servers that want independent accept queues or datagram sockets. It is not the same as SO_REUSEADDR, and enabling it does not automatically improve performance: the application still needs balanced workers, compatible socket setup, and measurements showing that listener contention is actually a bottleneck.
All sockets in a reuseport group must set the option before calling bind(). For sockets bound to the same address, Linux also requires the processes to have the same effective user ID, which helps prevent another local user from claiming the port. The exact address matters: a wildcard address, loopback address, and a specific interface address are not interchangeable assumptions. Confirm which local endpoint the service binds on each socket.
Set the option before binding
A minimal setup for a TCP listener has this order:
#define _GNU_SOURCE
#include <arpa/inet.h>
#include <netinet/in.h>
#include <stdint.h>
#include <sys/socket.h>
#include <unistd.h>
static int
CreateListener(uint16_t port)
{
struct sockaddr_in address = {0};
int enabled = 1;
int listener = socket(AF_INET, SOCK_STREAM, 0);
if (listener == -1)
return -1;
address.sin_family = AF_INET;
address.sin_addr.s_addr = htonl(INADDR_ANY);
address.sin_port = htons(port);
if (setsockopt(listener, SOL_SOCKET, SO_REUSEPORT,
&enabled, sizeof(enabled)) == -1
|| bind(listener, (struct sockaddr *)&address,
sizeof(address)) == -1
|| listen(listener, SOMAXCONN) == -1) {
close(listener);
return -1;
}
return listener;
}
The first listener and every additional listener must follow the same ordering and use compatible address families, transport protocols, and local addresses. Set the option before any bind attempt; setting it after a socket is already bound does not retroactively make the existing bind part of a reuseport group. This example binds an IPv4 wildcard address and checks the normal setup calls; production code should also configure close-on-exec and nonblocking mode as required, and report each failure with the socket’s intended address and worker identity.
For UDP, use SOCK_DGRAM and bind each socket to the same endpoint before receiving. The option is useful for both TCP and UDP, but their work models differ: a TCP listener receives new connection attempts that become established sockets, while a UDP socket receives datagrams and must preserve whatever packet-processing state the application requires. Do not assume that balancing listener sockets balances all work after accept or receive.
Distinguish reuseport from reuseaddr
SO_REUSEADDR and SO_REUSEPORT solve different binding problems. SO_REUSEADDR has platform- and protocol-specific semantics commonly used during restart or with certain address/port combinations. SO_REUSEPORT explicitly creates a group of sockets that can share an identical local address for traffic distribution, subject to Linux’s effective-UID restriction. One option is not a drop-in substitute for the other.
Many server examples set both options, but add each only for a documented need. A process that sets both without understanding the bind rules can obscure why a bind succeeded or failed. Test the intended combination on the target Linux kernel, address family, protocol, and deployment model. Also confirm that no unrelated service already owns a conflicting endpoint.
Understand default distribution and its limits
When a group is created, the kernel selects a member socket for incoming traffic using its reuseport selection mechanism. That spreads connections or datagrams across listeners without requiring every worker to contend on one shared listening socket. It does not promise perfectly equal request counts: connection lifetimes, traffic patterns, client address diversity, worker readiness, and application processing time can all create imbalance even if new flows are distributed evenly.
For TCP, a newly selected listener still has to accept promptly and process the connection. A worker that is slow, paused, or unable to accept can accumulate work differently from its peers. For UDP, a datagram is delivered to a selected socket, but the application must account for socket buffer drops, queue depth, packet ordering requirements, and state sharding. Flow-related selection can keep traffic associated with a flow on one socket, but client port changes, address changes, NAT, and selector behavior can affect that association.
Measure per-listener accepted connections or received datagrams, queue pressure, drops, CPU utilization, and request latency. Aggregate process-level throughput can hide one overloaded worker. Use a consistent workload and test both new-connection bursts and sustained traffic. If every listener is already idle and accept latency is negligible, SO_REUSEPORT may add operational complexity without improving the service.
Use a BPF selector only when needed
Linux can attach classic BPF or extended BPF programs using SO_ATTACH_REUSEPORT_CBPF or SO_ATTACH_REUSEPORT_EBPF to control selection within a reuseport group. This is useful when default selection is insufficient and the application has a concrete policy, such as steering based on packet metadata or worker availability. The program returns a socket index in the group; according to the socket API documentation, an invalid index falls back to the ordinary reuseport selection mechanism.
The socket API describes a BPF_PROG_TYPE_SK_REUSEPORT program for the extended-BPF path. Kernel support, verifier rules, available context, privileges, and attach behavior depend on the running kernel and program type. Verify the target kernel’s BPF documentation and headers, and load only a program whose selection logic is testable. BPF is not required for ordinary multi-listener scaling.
Socket membership can change over time. New sockets added to a reuseport group inherit its BPF program, and when a socket is removed the group can move another socket into its position. Therefore, do not treat socket indexes as permanent worker identities. A selector that caches an index must be updated with membership changes or can direct packets to an unintended worker. Keep a map from stable worker identity to current group index if the policy truly needs identity-aware selection.
Treat the selector as part of the service’s routing configuration. Validate its fallback behavior, worker restart sequence, program replacement, and behavior when a selected worker is overloaded. Log selection-policy version and listener membership without logging packet contents unnecessarily. Keep a known-good default selection path so a BPF policy failure does not strand new traffic.
Plan worker creation and restart carefully
Multiple processes can create their own listeners if credentials, addresses, and options match. Start workers in a controlled order, verify that every bind succeeds, and expose readiness only after the socket has joined the group and the worker can accept or receive. During rolling restart, test the interaction between old and new worker groups; a successful second bind does not prove that the old process will shut down without loss.
Set close-on-exec explicitly or use an atomic socket-creation flag where available so child processes do not unintentionally inherit listeners. Coordinate socket lifetime with the service manager and graceful shutdown logic. On shutdown, stop accepting new work, drain or close established connections according to the protocol, and then close the listener. UDP services need their own in-flight packet and state-drain policy because there is no connection close handshake for each datagram.
Containers and network namespaces add another boundary. A reuseport group exists for sockets that share the same network namespace and local endpoint; identical addresses in separate network namespaces are separate networking contexts. Check the namespace of each process, the effective UID inside the relevant credential context, and the bound address in that namespace. Do not debug a host listener when the service actually binds inside a container namespace.
Diagnose bind and balancing failures
If a later bind fails with EADDRINUSE, confirm every socket set SO_REUSEPORT before bind and that the address, protocol, namespace, and effective UID satisfy the group’s rules. Inspect existing listeners with tools such as ss in the same network namespace. A stale process, different address family behavior, or one worker binding a wildcard while another binds a specific address can explain a mismatch.
If all binds succeed but one worker receives little traffic, first verify that all sockets are actually receiving and that workers are ready. Compare listener membership, traffic tuple diversity, packet or connection rates, and per-worker metrics. A small number of long-lived TCP connections naturally limits how evenly work can be divided by connection selection. The application may need to parallelize after accept rather than create more listening sockets.
If BPF selection behaves unexpectedly, inspect program load and attach results, verifier logs, group membership changes, and the returned index bounds. Confirm the program type and context match the target socket option. Test invalid selection and removal/re-addition of a listener. Avoid changing the selector and worker topology simultaneously; otherwise a distribution shift is hard to attribute.
Benchmark the architecture rather than the option
Compare a single shared listener, a reuseport group, and any proposed BPF selector under the same hardware, kernel, workload, and worker count. Measure connection setup, throughput, CPU cost, tail latency, drops, and shutdown behavior. Include the production traffic shape: many short connections, a few long connections, or skewed UDP flows can produce different outcomes.
Test overload and failure. Pause one worker, restart one process, remove a socket from the group, and confirm that remaining workers handle traffic according to the service’s policy. Use a staging environment for packet-rate tests, and avoid using synthetic uniform flows as the sole proof of production balance. A lower accept-lock contention metric is not enough if an overloaded worker or application queue becomes the new bottleneck.
SO_REUSEPORT provides a kernel-managed group of compatible listeners, not an end-to-end load balancer. Set it before bind on every socket, verify the address and credential rules, measure actual per-worker distribution, and introduce BPF selection only when a tested policy needs it.
Related:
- Building an Isolated Network Namespace for Testing
- Linux Policy Routing with ip rule: Source-Based Paths and Verification
Sources: