Linux MSG_ZEROCOPY: Completion Ranges, Buffer Ownership, and Honest Benchmarks
Implement Linux MSG_ZEROCOPY as an asynchronous buffer-lifetime contract with error-queue completions, fallback handling, and workload benchmarks.
MSG_ZEROCOPY asks Linux to avoid copying user data into kernel-owned socket buffers for supported send paths. It does not promise that every send is physically copy-free, and it changes an important application invariant: the source memory cannot be safely modified or reused until the kernel reports completion. The feature trades memory-copy CPU cost for page pinning, accounting, and asynchronous completion handling. For small writes, that overhead can be higher than copying.
Adopting the flag is therefore a buffer-lifecycle change, not a one-line optimization. A correct program enables the socket option, tags eligible sends, drains completion notifications from the error queue, handles completion ID ranges, and benchmarks against ordinary sends under realistic transport conditions.
Supported path and opt-in
Linux documents MSG_ZEROCOPY for TCP, UDP, and VSOCK with the virtio transport. Availability still depends on the running kernel and socket path. The documented interface requires the application to set SO_ZEROCOPY on the socket before using the message flag. This explicit opt-in prevents applications that accidentally pass an undefined flag from silently changing behavior.
For eligible sends, the application can pass MSG_ZEROCOPY to send(), sendto(), sendmsg(), or sendmmsg(). Normal copied sends can be mixed with copy-avoidance sends on the same socket. The return value reports bytes accepted by the socket operation, not that the network peer received them and not that the source buffer can be reused.
The kernel may copy data even when zero-copy was requested. The documented loopback path is an example: local TCP/UDP delivery can defer a copy so the sender is not held hostage to a local receiver that does not read. Notifications can report that a copy occurred. A test over loopback or a local virtual device can therefore show no benefit and should not be generalized to a remote NIC workload.
Buffer ownership and completion notifications
With a conventional copied send, the application can typically reuse the source buffer after the call returns, subject to the API’s normal rules. With MSG_ZEROCOPY, the kernel may retain references to user pages while network transmission is in flight. The application must not modify those bytes until it receives the corresponding completion notification. Modifying them early may alter data that is still being transmitted, corrupting the application’s own stream.
The kernel reports completions through the socket error queue, read with recvmsg() and MSG_ERRQUEUE. These are not ordinary network errors. The control message includes a sock_extended_err record with an origin identifying zero-copy completion and an ID interval describing one or more completed sends. Applications need to map that interval to buffers and release only the buffers whose sends have completed.
Completion IDs are sequence numbers associated with zero-copy sends and can wrap. One notification can cover a range because completions may be coalesced. Code must correctly handle a range that crosses the counter wrap point, multiple outstanding buffers, partial sends, and interleaved copied sends. Do not assume one completion per send() or that notification order matches a simple one-buffer/one-callback pattern.
The extended error’s code can indicate that the kernel copied the data rather than completing a zero-copy operation. That still releases the buffer according to the completion contract, but it changes the performance interpretation. Keep the completion path active even if a particular benchmark rarely reports fallback copies.
Event loop design
An event loop that uses zero-copy should monitor both writable readiness and error-queue readiness. When a send is accepted, record the buffer and its zero-copy ID association. When an error-queue notification arrives, validate the control-message level, type, length, origin, and completion range before releasing memory. A malformed or truncated control message must not cause the program to free arbitrary buffers. If the error queue is not drained, outstanding pages can remain pinned and new sends may fail due to socket memory or locked-page accounting limits.
Backpressure must be explicit. Bound the number and total size of in-flight buffers. When the limit is reached, stop submitting new data until completions free capacity. If a send returns a short count, account for the accepted prefix and retry the remainder with correct IDs and buffer lifetime; do not mark the entire source allocation reusable because one syscall returned.
Handle socket teardown carefully. Closing the socket is not a substitute for a correct in-flight ownership model when the application still has references to shared buffers or uses a memory pool. Define how shutdown drains completions, abandons outstanding sends, and prevents a buffer from being returned to another producer prematurely. If the process crashes, the kernel will reclaim resources, but application-level shared memory, DMA, or cross-process ownership may require separate coordination.
Errors and limits
The kernel documentation describes failures such as ENOBUFS when socket option memory or locked-page limits are exceeded. A burst of sends can therefore fail even though the ordinary copied path worked. Record the socket’s error, in-flight byte count, completion queue health, and system resource limits before changing global limits. Raising limits without bounding application concurrency can convert a visible send failure into excessive pinning and memory pressure.
Some devices and paths cannot use the optimization efficiently or at all. Fragmentation, alignment, protocol segmentation, memory pressure, and driver behavior can affect the path. The API is an opportunity for copy avoidance, not an end-to-end guarantee that NIC DMA reads the exact original user pages.
The typical break-even size is workload-dependent. Kernel documentation notes that page pinning and completion overhead often make the feature most useful for larger writes, with an approximate scale around 10 KB in the documented implementation. Treat that as a starting point for measurement, not a universal threshold. Coalescing, CPU architecture, transport, payload size distribution, and batching all matter.
Benchmark without misleading yourself
Compare the ordinary send path and zero-copy path using the same application workload, payload distribution, peer, transport, CPU placement, and network configuration. Measure CPU cycles per useful byte, throughput, tail latency, page-pin/completion overhead, retransmissions, and memory footprint. Separate warm-up from steady state and run enough repetitions to account for frequency scaling and other host noise.
Use a remote host for representative testing when production traffic is remote. Local loopback intentionally differs. A veth pair across network namespaces may still exercise local delivery behavior described by the kernel documentation and is not an adequate proxy for a physical NIC without validating the actual path. Verify payload integrity at the receiver, and include a slow receiver to test backpressure rather than benchmarking only an always-draining sink.
Instrument how often completion records indicate copied fallback, how many buffers are in flight, the maximum completion delay, and the age of the oldest unreleased buffer. Ensure benchmarks include both small control messages and large bulk writes if the application sends both. A result that improves throughput only by pinning unbounded memory is not an operational win.
Rollout and acceptance
Deploy behind a feature flag or a per-socket threshold. Keep a copied-send fallback for unsupported kernels and for payloads below the measured crossover point. Log the kernel release, socket family, transport, enabled option, buffer size, completion behavior, and errno distribution. Roll back if completion processing stalls, locked pages rise unexpectedly, or tail latency worsens.
Before accepting the optimization, verify that every zero-copy send has a matching completion lifecycle, the error queue is drained under load, ID wrap and coalesced ranges are tested, short sends are handled, buffer mutation is prevented, and shutdown is bounded. Use fault injection or a focused test harness to exercise slow receivers and constrained memory in a non-production environment.
The correct abstraction is “send data and later learn when this source storage can be reused,” not “send without copying.” That model keeps memory ownership and transport completion explicit and avoids turning an apparent micro-optimization into silent data corruption or resource exhaustion.
Related:
- Linux splice: Pipe-Backed Transfers Without a Userspace Copy
- Linux SO_TIMESTAMPING: Trace Packet Time Through the Network Stack
Sources: