FreeBSD TCP Congestion Control: Selecting Algorithms and Measuring Real Paths
Operate FreeBSD's modular TCP congestion control with per-connection selection, current defaults, socket evidence, controlled benchmarks, and safe rollout criteria.
TCP congestion control governs how a sender probes for available network capacity and reacts to congestion signals. It is not the same as TCP flow control, which limits sending to the receiver’s advertised window. It is also not a traffic shaper, packet scheduler, or a repair for a broken link. Changing a congestion controller can alter throughput, queueing delay, fairness, and burst behavior, but it cannot make an undersized uplink, a lossy radio, or an overloaded peer disappear.
FreeBSD makes congestion control modular. The mod_cc(4) framework exposes algorithm names that can be built into the kernel or made available as modules. A system-wide default is selected for connections unless an application overrides it using the TCP TCP_CONGESTION socket option. The installed kernel and release determine what is available. In the FreeBSD 15.1 manuals, CUBIC is the default TCP congestion control algorithm; do not copy this assumption to a different release, custom kernel, jail, or TCP stack without checking the running host.
Establish the layer you are measuring
Before changing anything, write down the symptom in measurable terms. A transfer may plateau below the link rate, show high tail latency while a bulk flow runs, retransmit heavily, or perform differently in one direction. Those outcomes can come from the sender’s congestion window, the receiver’s advertised window, path MTU, loss, queueing, shaping, CPU cost, storage, application pacing, or remote endpoint limits.
Congestion control affects TCP send behavior in response to acknowledgments and congestion events. The FreeBSD mod_cc(9) interface describes events such as an explicit congestion notification, retransmission timeout, and duplicate acknowledgments being delivered to the algorithm. The algorithm updates per-connection state; it does not select the route or configure a NIC’s offloads. A result attributed to congestion control should therefore include both socket-level evidence and end-to-end measurements.
Start with a read-only snapshot:
freebsd-version -kru
sysctl net.inet.tcp.cc.available
sysctl net.inet.tcp.cc.algorithm
netstat -C -n
netstat -T -n
The available-algorithm MIB is read-only. The algorithm MIB reports the default and can change it when written. netstat -C displays congestion-control algorithm and diagnostic information for TCP sockets; netstat -T displays TCP control-block information including retransmission, out-of-order, and zero-window evidence. Save these outputs at the start and end of a test. A single counter snapshot cannot establish a rate, and aggregate retransmissions do not identify which network segment dropped packets.
Understand what an algorithm change does
The modular interface is useful because an application can select an algorithm for a particular connection instead of making every system TCP socket use the same policy. The name must be present in the running kernel’s available list. A socket option is a deliberate application-level choice, not a persistent kernel setting:
#include <sys/types.h>
#include <sys/socket.h>
#include <netinet/in.h>
#include <netinet/tcp.h>
const char algorithm[] = "cubic";
if (setsockopt(fd, IPPROTO_TCP, TCP_CONGESTION,
algorithm, sizeof(algorithm)) == -1) {
/* Handle errno; do not silently assume the request took effect. */
}
This example is a focused client-side fragment. Apply the option to the socket at a point allowed by the target release and application lifecycle, check the return value, and query the resulting socket where supported. sizeof(algorithm) includes the trailing NUL byte, which the FreeBSD implementation accepts as part of the name. An unavailable module, a typo, an unsupported algorithm, or an invalid option timing can fail; log that failure rather than quietly falling back while reporting that the experiment used the requested algorithm.
The system default can be read and, if intentionally changed, written through net.inet.tcp.cc.algorithm. Before a global change, inspect net.inet.tcp.cc.available and confirm the proposed name appears. Prefer an application-scoped experiment first: it limits the blast radius, creates a clear comparison group, and avoids changing unrelated services. A global default is a host-wide policy decision and may affect applications that never expected the new behavior.
The algorithm name is not necessarily the name of an entire TCP implementation. FreeBSD documents the tcp_bbr(4) implementation as a separate TCP function block selected through the TCP-functions mechanism. Do not assume that bbr can be selected by the mod_cc list or that changing net.inet.tcp.cc.algorithm enables it. Follow the exact manual for the implementation and inspect the relevant available list before testing. Mixing a TCP stack selection with a congestion-control module selection produces an invalid comparison.
Choose by hypothesis, not reputation
The FreeBSD cc_cubic(4) manual describes CUBIC as designed for fast and long-distance networks, with a congestion window that grows as a function of time since the last congestion event. cc_newreno(4) implements NewReno behavior and is useful as a conservative comparison point, though it is not the current 15.1 GENERIC default. Other modules have their own assumptions, code paths, and deployment requirements. A name or benchmark result from a different operating system does not guarantee the same implementation or benefit on FreeBSD.
Make a falsifiable hypothesis before selecting an alternative. For example: “Under a sustained bulk transfer across a high-bandwidth, high-round-trip path, the observed sender rate is constrained by loss recovery rather than CPU, receiver window, or storage; compare the current default with a named available alternative while measuring goodput and unloaded/loaded latency.” That hypothesis identifies the traffic shape, path, controls, and success metric. It does not presume which algorithm must win.
Congestion control has tradeoffs. A controller that achieves a higher bulk rate can increase queue occupancy or affect other flows. A delay-sensitive service may care more about loaded latency than peak throughput. A short-lived request may finish before a long-lived flow reaches its steady behavior. Competing senders, receiver buffering, ACK behavior, and middleboxes also alter the result. Measure with representative workloads and multiple runs, including a baseline with no competing test flow and a loaded case that resembles production.
Run a controlled comparison
Use the same endpoints, route, interface, MTU, application binary, socket buffer policy, data set, and test duration for each run. Record FreeBSD release and kernel build, NIC driver and firmware, selected algorithm, module availability, test direction, path RTT, and traffic shaping or VPN layers. Let the test run long enough to reveal startup and steady-state behavior, but keep it isolated from unrelated production traffic unless the test plan explicitly evaluates coexistence.
Measure more than bytes per second. Capture transfer goodput, completion time, retransmission changes, loaded and unloaded RTT, CPU consumption, and any application-visible timeouts. Use netstat -C -n to verify the algorithm attached to the sockets under test. netstat -T -n can help identify retransmit, out-of-order, and zero-window observations, but interpret those counters with endpoint and packet-path data. A zero receive window is a receiver/application backpressure symptom, not a signal that choosing a different sender congestion algorithm will solve it.
Run repetitions in alternating order where practical. If every baseline run occurs in the morning and every alternative run during a busy period, time-of-day load becomes a confounder. Preserve raw measurements, not only averages. Report median and tail values, variation, and test conditions. When multiple TCP connections are used, report whether each used the intended option; a single correctly configured socket does not prove a whole application selected it.
Safe deployment and rollback
First establish that the candidate is present on every target kernel and survives the intended boot or module lifecycle. A module loaded on one test host is not a production rollout plan. If an application selects its algorithm per socket, deploy behind a feature flag or service-specific configuration and preserve a way to turn it off. For a global default, stage on one canary, record the previous default, and test all critical flows including management, replication, monitoring, and backup traffic.
Define acceptance criteria before deployment: no regression beyond agreed throughput or latency bounds, no unexplained rise in retransmissions or timeouts, stable CPU and memory use, and a confirmed rollback command or configuration. Compare active sockets after rollout instead of assuming existing connections changed when a default was updated. New connections and long-lived sessions can have different histories; drain or reconnect them only through the service’s normal maintenance procedure.
Do not stack unrelated TCP sysctl tuning on the same change. Receive-window autotuning, send buffers, delayed acknowledgments, segmentation offload, path MTU, and congestion control affect distinct parts of the path. Change one dimension, keep the test reproducible, and revert if results are ambiguous. Avoid undocumented values copied from a Linux guide: names, defaults, algorithm implementations, and controls differ across systems and releases.
Evidence that supports a conclusion
A useful incident record contains the exact kernel version, algorithm availability and default, per-socket algorithm evidence, socket diagnostics, endpoint and path description, workload parameters, repeated throughput and latency measurements, and any changes made. State which hypotheses were ruled out and which remain uncertain. For example, if the receiver advertises a zero window, the next investigation belongs at the receiver or application, not in the congestion-control algorithm. If the connection uses a different TCP stack than expected, do not interpret algorithm-only documentation as the full behavior.
The goal is not to find a universal best controller. It is to verify which implementation actually ran, demonstrate a repeatable benefit for the target workload, and preserve acceptable fairness and latency for other traffic. FreeBSD’s modular design enables precise experiments, but only careful measurement turns that flexibility into a production improvement.
Related:
- FreeBSD Networking Internals: Interfaces, Routing, and netstat
- FreeBSD BPF Packet Capture: Filters, Buffers, and Loss-Aware Analysis
Sources: