Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux splice: Pipe-Backed Transfers Without a Userspace Copy

Understand Linux splice's pipe requirement, offsets, short transfers, nonblocking limits, and how to preserve buffered bytes across backpressure.

Linux splice() moves data between two file descriptors while avoiding the normal copy through a userspace buffer. At least one endpoint must be a pipe. That constraint is fundamental: a pipe supplies kernel-managed buffer slots that can hold references to pages while data moves between a file, socket, or another pipe.

The phrase “zero-copy” needs care. splice() can avoid copying bytes from the kernel into application memory and back, and the kernel may pass references to pages between pipe buffers. It does not guarantee that no data copy occurs internally, that every filesystem supports the transfer, or that the operation is faster for every workload. The API changes where the data path runs; it does not remove ownership, backpressure, partial progress, or descriptor-lifetime concerns.

The pipe is part of the data path

A typical file-to-socket relay uses two calls. The first moves bytes from a regular file into the write end of a pipe; the second moves bytes from the read end to a socket. The pipe is not optional staging ceremony. It is the required kernel buffer used by the interface.

off_t source_offset = 0;
ssize_t staged = splice(file_fd, &source_offset,
                        pipe_write_fd, NULL, 64 * 1024,
                        SPLICE_F_NONBLOCK);
if (staged > 0) {
    /* Forward up to staged bytes from pipe_read_fd to the socket. */
    ssize_t sent = splice(pipe_read_fd, NULL, socket_fd, NULL,
                          (size_t)staged, SPLICE_F_NONBLOCK);
    /* If sent is short, the unsent bytes remain queued in the pipe. */
}

This fragment shows the data-path shape, not a complete pump. A production loop must remember any bytes still buffered in the pipe and forward them before staging more input. The source offset advances when bytes are moved into the pipe, while the socket may accept fewer bytes than that amount. If the program loses track of this difference, it can skip data, duplicate it, or incorrectly report the transfer complete.

The pipe has finite capacity, which depends on kernel configuration and resource limits. Treat it as backpressure storage, not an unbounded queue. A full pipe prevents more input from being staged until the consumer drains it. A nonblocking design should register the relevant descriptors with poll or epoll and resume when the operation can make progress.

Offset pointers describe seekable endpoints

For a pipe endpoint, its corresponding offset pointer must be NULL. For a seekable non-pipe endpoint, a NULL pointer uses and advances the open file description’s current offset; a non-NULL pointer supplies an explicit offset and is updated without changing that shared file position.

Use explicit offsets when the transfer should not mutate a shared file offset or when the application needs a separate progress record. An offset is not valid for a pipe or another nonseekable endpoint. Passing one can produce ESPIPE or EINVAL, depending on the condition and interface path.

If both endpoints are pipes, both offset pointers are NULL. Linux has permitted pipe-to-pipe splice since kernel 2.6.31; code targeting older kernels must treat that as a compatibility boundary. Passing the same pipe as both endpoints is invalid. Validate that descriptor roles are distinct and correctly opened before entering the loop.

A short result is normal progress

splice() transfers up to the requested size, not necessarily the entire request. A successful result can be shorter because the pipe has limited room, the socket has limited capacity, the source currently has fewer available bytes, or the filesystem and endpoint make less progress than requested. Track the exact returned byte count and retry only the remaining transfer.

A return value of zero means end of input for the operation’s source. When the source is a pipe, zero means there is no data to transfer and no writers remain connected. It is not a generic “temporarily nothing available” result to spin on; with nonblocking descriptors, lack of immediate progress is normally reported through EAGAIN.

EINTR means the call was interrupted before completing its reported transfer. A robust event loop retries after interruption while preserving state, but it should not rebuild offsets from stale values or reset the pipe accounting. A successful positive return is progress and must be recorded before the next syscall.

Nonblocking behavior is not automatic for every endpoint

SPLICE_F_NONBLOCK makes the pipe operations nonblocking. It does not by itself guarantee that the non-pipe endpoint cannot block; the file or socket may still block unless its open file description has O_NONBLOCK. If the goal is a never-blocking event loop, configure nonblocking behavior on the relevant descriptors too, and confirm the semantics for each endpoint type.

When a nonblocking call returns EAGAIN, wait for the appropriate readiness event instead of spinning. The direction matters: a pipe may be readable while a socket is not writable, or the pipe may be full while the source is ready. An epoll state machine should track whether it needs source readability, pipe room, or destination writability and should drain until the next operation would block when edge-triggered readiness is used.

The SPLICE_F_MORE flag is a hint that more data will follow in a later splice, useful when the output is a socket. It does not guarantee packet boundaries or application framing. SPLICE_F_MOVE is only a hint to move pages and has been a no-op in Linux since 2.6.21; do not make correctness or performance claims depend on it. SPLICE_F_GIFT is not used by splice().

Model a relay as an explicit state machine

A file-to-socket relay should maintain at least these facts: the source offset, the number of bytes currently buffered in the pipe, whether source EOF has been observed, whether the destination has failed, and whether the connection should be closed after pending data drains. If a splice into the pipe stages 32 KiB and the socket accepts only 12 KiB, the remaining 20 KiB are already in the pipe and must be sent before more source data is read.

Use one owner for each pipe end and close every unused duplicate after fork(). A process that accidentally retains a pipe write end can prevent the reader from observing EOF. A leaked read end can also change broken-pipe behavior. O_CLOEXEC on pipe creation and deliberate descriptor inheritance reduce these lifecycle bugs.

Do not hold application locks while a potentially blocking splice waits for pipe or socket progress. A blocked transfer can prevent another thread from draining the pipe or changing the state that would make the operation safe. In a multithreaded design, assign a clear owner to each transfer state or protect the state with short critical sections around bookkeeping rather than around I/O waits.

Compatibility and fallback are endpoint-specific

splice() is Linux-specific and the target filesystem or descriptor pair may not support it. Unsupported combinations can fail with EINVAL or other errors, and a socket, filesystem, or device can impose its own constraints. A fallback through read() and write() is reasonable when the application can tolerate userspace buffering, but it must start from the exact byte boundary that remains after any successful splice operations.

Never interpret every EINVAL as “try buffered I/O.” It can indicate an invalid offset, a non-pipe pair, append mode, or the same pipe used as both endpoints. Validate the programming contract first. Likewise, EPIPE, ECONNRESET, ENOSPC, and EIO are meaningful destination or source failures that should not be hidden behind a silent fallback.

The kernel may avoid copying data into userspace, yet an application still needs checksums, encryption, compression, content inspection, or transformation in many workflows. Those requirements may force bytes through userspace or another kernel API. Select splice() only when the required semantics fit the interface; a benchmark that omits error handling and backpressure is not evidence that the production path is superior.

Verify data, lifecycle, and pressure

Test short pipe capacity, a deliberately slow receiver, disconnects during transfer, signal interruption, nonblocking EAGAIN, source EOF, filesystems that reject the operation, and a fallback after partial progress. Compare byte counts and checksums against a buffered reference path. Include cases where both endpoints are pipes and ensure offsets are NULL.

Observe bytes staged separately from bytes delivered. Record pipe occupancy where available, time spent waiting for source and destination readiness, fallback counts, and failures by errno. Validate descriptor closure and child inheritance, especially in fork/exec pipelines. Make shutdown drain or discard buffered pipe bytes according to an explicit policy.

splice() can simplify a kernel-resident data path, but the pipe is still a finite queue and the application still owns the transfer protocol. Track exact progress, respect endpoint offset rules, configure nonblocking behavior deliberately, and treat “zero-copy” as an implementation opportunity rather than an unconditional performance guarantee.

Related:

Sources:

Comments