Skip to content
Tech HistoryDeep Dive Published Updated 8 min readViews unavailable

TCP Congestion Control: From the 1988 Collapse to CUBIC

A sourced history of TCP congestion control, from the 1980s collapse and Reno to SACK, NewReno, and modern CUBIC behavior.

TCP congestion control grew out of a practical failure: networks could have usable link capacity and still deliver far less useful data when senders reacted badly to packet loss and queues. In their 1988 paper, Van Jacobson and Michael Karels described congestion collapse and a set of algorithms grounded in packet conservation. In simplified terms, a sender should not inject new data faster than acknowledgments and network feedback justify. If it does, retransmissions can consume capacity without increasing useful delivery.

The paper’s contribution was not one magic congestion-window formula. It joined measurements, retransmission behavior, and sender rate control into a feedback system. That system became the basis for the familiar TCP mechanisms: slow start, congestion avoidance, fast retransmit, and fast recovery. Each answers a different question: how to probe an unfamiliar path, how to increase use cautiously, how to infer a likely missing segment before a timer expires, and how to continue after loss without immediately restarting from the smallest window.

Congestion is not receiver flow control

TCP has two distinct limits on outstanding data. The receiver advertises rwnd so a sender does not overrun the receiving application’s buffer. The sender maintains cwnd, a congestion-control limit intended to regulate how much unacknowledged data is in the network. At a given moment, the lower of those two limits governs transmission. The receiver can therefore be ready for more bytes while the path is congested, or the path can have room while the receiver’s buffer is the tighter constraint.

The RFC terminology also distinguishes the congestion window from FlightSize: data already sent but not yet cumulatively acknowledged. That distinction matters during loss response. cwnd is a sending limit; FlightSize describes outstanding traffic. Substituting one for the other in a congestion calculation can produce the wrong response, especially when receiver flow control has kept actual in-flight data below the sender’s nominal window.

The sender’s maximum segment size (SMSS) is the payload size it can send in a segment under the connection’s path and negotiated constraints. Window sizes may be represented in bytes or segment units, but analyses must use consistent units. A rough bandwidth-delay product illustrates why a window matters: a 100 Mbit/s path with a 40 ms round-trip time needs about 4 Mbit, or 500,000 bytes, in flight to keep that idealized path continuously occupied. Headers, acknowledgments, competing traffic, protocol overhead, and actual path behavior make this only an estimate, not a target that guarantees performance.

The four classic mechanisms

Slow start is a probe, not a promise to send slowly forever. A connection starts with a limited congestion window and increases it as new acknowledgments arrive. Under the textbook case of full-sized segments and prompt acknowledgments, that ACK-driven growth can approximately double the window per round-trip time. The actual increase is tied to acknowledged data, so delayed ACKs, application limits, receiver limits, and packetization affect the result. Growth stops when the window reaches the slow-start threshold (ssthresh) or congestion is detected.

Congestion avoidance grows the window more cautiously. The RFC 5681 Reno-style rule is roughly one SMSS per round-trip time, often described as additive increase. When loss signals congestion, a sender reduces its operating window, the multiplicative-decrease part of the classic AIMD pattern. The purpose is not to punish one packet loss; it is to back away from a sending rate that may be contributing to an overloaded bottleneck, then cautiously test capacity again.

Fast retransmit uses duplicate acknowledgments as an early loss clue. In the classic algorithm, three qualifying duplicate ACKs cause the sender to retransmit the segment beginning at the oldest unacknowledged sequence number without waiting for the retransmission timer. Duplicate ACKs can arise from loss, but also from reordering or replication, so the count is a practical inference rather than proof that a router dropped a packet.

Fast recovery governs the sending window after that retransmission. Duplicate ACKs indicate that later segments are still reaching the receiver, so the sender has evidence that some traffic continues to leave the network and the ACK clock has not entirely stopped. Reno-style recovery reduces the window while allowing carefully limited progress, instead of dropping straight back into slow start for every fast-retransmit event. The exact state transitions are defined in the RFC; the useful historical point is that recovery preserves information about packets still moving through the path.

An RTO timeout is a different signal. It means the sender did not receive the expected acknowledgment in time, so the classic algorithm backs off more sharply and restarts with a loss window of one full-sized segment, then slow-starts toward the reduced threshold. This is why a timeout can cost substantially more throughput than a fast retransmit. It is also why packet captures and counters should distinguish duplicate-ACK recovery from retransmission-timeout recovery instead of treating all retransmissions as equivalent.

At a high level, the old Reno-style progression can be sketched as follows. This is explanatory pseudocode, not a replacement for the normative rules and edge cases in the RFCs:

on new acknowledgment:
    if cwnd < ssthresh:
        grow cwnd using slow-start rules
    else:
        grow cwnd approximately one segment per RTT

on three qualifying duplicate acknowledgments:
    reduce ssthresh using outstanding FlightSize
    retransmit the earliest missing segment
    enter fast recovery

on retransmission timeout:
    reduce ssthresh using outstanding FlightSize
    set cwnd to one full-sized segment
    retransmit and resume with slow start

From Jacobson and Karels to RFCs

RFC 2001 documented slow start, congestion avoidance, fast retransmit, and fast recovery for interoperable TCP implementations. RFC 2581 later revised that description, and RFC 5681, published in 2009, consolidated the Reno-style requirements and additional behavior such as restarting after an idle period and ACK generation considerations. This progression converted a response to observed Internet failure into shared, testable protocol guidance.

The classic algorithms did not solve every loss pattern equally well. Cumulative acknowledgments normally report the next byte expected, which can leave a sender uncertain about which later data arrived after multiple losses. The Selective Acknowledgment (SACK) option, standardized in RFC 2018, lets a receiver report non-contiguous blocks of data it has already received, provided the peers negotiated its use. That extra information can help a sender recover several losses in one flight without retransmitting data it knows arrived.

NewReno, specified in RFC 6582, improves fast recovery when SACK is unavailable. It uses a partial acknowledgment during recovery as evidence that another segment from the same flight may still be missing, and continues recovery accordingly. SACK and NewReno are related extensions to loss recovery, but they are not synonyms: one adds acknowledgment information, while the other changes how a sender interprets partial acknowledgments during recovery.

The congestion-avoidance growth function has also evolved. CUBIC is specified in RFC 9438 for fast and long-distance networks. Rather than use only Reno’s linear additive increase, CUBIC shapes its congestion-window growth with a cubic function based on elapsed time since the last congestion event. RFC 9438 describes how it retains ACK-clocked operation and integrates with loss-recovery mechanisms, while choosing a different increase curve and decrease factor. Consequently, RFC 5681 remains an important baseline for the classic algorithms, but it is not accurate to imply that every modern TCP connection follows only the original Reno growth rule. The negotiated features, host operating system, selected congestion-control algorithm, and network signals all matter.

Explicit Congestion Notification (ECN) is another important qualification to a loss-only account. When enabled and supported along the path and by both endpoints, routers can mark packets to signal congestion rather than relying exclusively on packet loss. The TCP congestion response still needs to reduce load; ECN changes how the signal can be conveyed, not the need for a feedback loop.

How to use the history when diagnosing a real path

Congestion control explains why packet loss, delay, retransmission timers, receiver acknowledgments, and queue behavior cannot be debugged independently. A sender that retransmits too aggressively can add load to a congested path; one that waits too long can leave available capacity unused. A low application throughput reading alone does not identify which case applies.

For a controlled investigation, compare several observations over the same interval: application goodput, RTT and its variation, retransmitted segments, duplicate ACKs, timeouts, receiver-window limitation, interface or NIC drops, and the congestion-control algorithm reported by the sending host. Check whether the application continuously has data ready to send; a small or bursty workload may never fill the congestion window. Check also whether the receiver window, a host pacing limit, CPU saturation, or a middlebox is the actual bottleneck. A single retransmission counter without a denominator or time series can mislead.

When studying a historical statement, keep three dates separate: when a problem was observed, when a paper proposed a remedy, and when a standards document specified behavior. The 1988 paper is the source for the original intervention; the RFCs define particular protocol requirements and later variants. The broader lesson is disciplined feedback: measure what the path can sustain, react to congestion signals, and avoid turning recovery into another source of overload.

Related:

Sources:

Comments