Linux copy_file_range: Efficient File Copies with Explicit Boundaries
Use copy_file_range for kernel-assisted copies while handling partial progress, filesystem capability, offsets, sparse files, and safe fallback paths.
Copying a file through a userspace buffer usually means reading bytes into application memory and then writing them back to the kernel. Linux copy_file_range() can move a range between two file descriptors without routing the bytes through a userspace buffer. A filesystem may use this opportunity for an efficient server-side copy, a reflink-like operation, or another implementation-specific optimization.
The syscall is an optimization opportunity, not a promise of zero physical copying, atomicity, metadata preservation, or identical behavior on every filesystem. Reliable software treats it as one transfer mechanism inside a carefully specified copy workflow.
Define the copy contract first
Before choosing an API, define whether the operation copies the entire source or a byte range, whether source and destination can change concurrently, whether sparse-file layout matters, which metadata must be preserved, and what should happen after interruption. copy_file_range() handles bytes in a range. It does not copy owner, mode, timestamps, ACLs, extended attributes, capabilities, file flags, or directory entries.
Open the source for reading and destination for writing under a trusted path policy. Use O_CLOEXEC where descriptor inheritance is not intended. If untrusted path components are possible, directory descriptors and constrained path resolution may be necessary before this syscall is reached. The function operates on already-open descriptors; it does not validate that those descriptors identify the paths or files your higher-level policy intended.
The destination must not be opened with O_APPEND; copy_file_range() rejects that open-file-description mode because it conflicts with the operation’s offset contract. The source and destination must be regular files, and both descriptors must have the required access mode.
The source and destination may refer to the same file only when the requested ranges do not overlap. Overlapping ranges in one file are rejected. If another process can modify either file during the copy, define whether the result may contain a mixture of versions; this syscall is not a snapshot mechanism. A consistent backup needs an application-level snapshot, filesystem snapshot, locking scheme, or source that is immutable for the duration.
Use explicit offsets when the file positions are shared state
When an offset pointer is NULL, the syscall uses and advances that open file description’s current file offset. When the pointer is non-NULL, the pointed-to offset is used and updated while the descriptor’s shared file position is left unchanged. This distinction matters if descriptors are duplicated or shared with other threads, because the open file description’s offset can be shared too.
#define _GNU_SOURCE
#define _FILE_OFFSET_BITS 64
#include <errno.h>
#include <stdint.h>
#include <sys/types.h>
#include <unistd.h>
int copy_range_loop(int source_fd, int destination_fd,
off_t *source_offset, off_t *destination_offset,
uint64_t remaining) {
const size_t chunk_limit = 1024 * 1024;
while (remaining > 0) {
size_t request = remaining < chunk_limit
? (size_t)remaining : chunk_limit;
ssize_t copied = copy_file_range(source_fd, source_offset,
destination_fd, destination_offset,
request, 0);
if (copied > 0) {
remaining -= (uint64_t)copied;
continue;
}
if (copied == 0)
return 1; /* Source reached EOF before the requested range ended. */
if (errno == EINTR)
continue;
return -1;
}
return 0;
}
The return values here are an example contract: zero means all requested bytes copied, one means the source ended early, and minus one means an error. Production code should attach structured error details and consider whether a short source is expected. The function deliberately accepts explicit offsets and a bounded request size. Its caller must validate that offsets and remaining fit the platform’s supported file-offset range.
A successful call may copy fewer bytes than requested. Continue from the updated offsets until the defined range is complete or a short-source policy is reached. A zero result means the input is at or beyond end-of-file, not “retry until something appears.” Retrying zero without changing state creates an infinite loop.
Understand the filesystem and kernel capability boundary
The syscall dates to Linux 4.5, and its implementation was substantially reworked in Linux 5.3. On Linux 5.19 and newer, cross-filesystem copies can work when both filesystems are the same type and implement the operation; this behavior has also been backported to stable kernels. The Linux man page warns that the earlier 5.3–5.18 cross-filesystem fallback could report success without copying on some virtual filesystems. Do not encode “cross-filesystem always returns EXDEV” or “cross-filesystem always works” as a universal rule, especially on older kernels.
The filesystem may implement a server-side copy or copy-on-write sharing. These can reduce data movement, but the result still has ordinary file semantics and can have filesystem-specific performance and resource behavior. On sparse input, a copy can expand holes into allocated zero-filled data. If sparsity is part of the contract, inspect holes with SEEK_DATA and SEEK_HOLE where supported and copy data extents deliberately, preserving holes by advancing destination offsets rather than writing zero blocks.
An unsupported-operation error is a capability result, not proof that the source or destination is corrupt. Conversely, a successful return does not prove that the filesystem used a zero-copy or reflink path. Measure the target workload and verify actual destination content and metadata independently.
Implement a safe fallback without losing progress
Applications often need a portable read/write fallback when the kernel reports that the optimized operation is unavailable for the file pair. The fallback must preserve the bytes already copied, continue at the exact source and destination offsets, handle short reads and writes, retry interrupted operations according to their return values, and avoid an infinite loop when a write returns zero.
Do not respond to every copy_file_range() error by blindly restarting with a userspace copy from offset zero. If the syscall copied earlier chunks before a later failure, restarting can duplicate data or overwrite the wrong region. Keep a transfer ledger: bytes completed, current offsets, whether truncation is allowed, and which metadata stage has completed. Use fallback only for errors your policy classifies as “optimized path unavailable.” Errors such as ENOSPC, EIO, EFBIG, permission failures, and immutable destination state are not generic signals to retry another mechanism without analysis.
The kernel’s cross-filesystem behavior is version and filesystem dependent. A caller may elect to fall back on EXDEV for a generic copy, but it must also consider EOPNOTSUPP and ENOSYS where relevant and must not hide unexpected filesystem errors. If the source or destination is a network filesystem, define what the reported success means under that filesystem’s durability and caching contract. A successful copy is not automatically an fsync() durability guarantee.
For a robust replacement operation, a common pattern is to create a temporary file in the destination directory, copy into it, apply required metadata, flush the file according to the durability requirement, and atomically rename it into place. When crash durability of the name change matters, also sync the containing directory after the rename where the filesystem supports that contract. This avoids exposing a partially written destination under the final name, but atomic rename and durable rename are different claims; the exact sequence depends on directory permissions and filesystem semantics.
Treat metadata, truncation, and destination state separately
copy_file_range() writes into the requested target range and overwrites bytes already there. It does not automatically truncate the destination to the new logical size. If the destination was longer than the copied range, a stale suffix can remain. Decide whether the operation should preserve untouched destination bytes, truncate to an exact length, or write into a fresh temporary file.
When truncating or replacing, ensure there is no concurrent writer that can race the copy. When preserving permissions and ownership, remember that applying metadata before writing may have security implications, and applying it afterward may fail. Set restrictive creation permissions initially, copy bytes, then apply the metadata allowlist that the application is authorized to preserve. Do not copy setuid bits, file capabilities, ACLs, or security labels by accident.
The syscall also does not calculate a checksum. If integrity verification is required, compute and compare hashes using a defined source snapshot or immutable input. A checksum calculated after copying from a source that changed concurrently cannot establish that the destination matches any single source version.
Test the real matrix
Exercise regular files, empty files, exact EOF, a source shorter than requested, partial progress, explicit offsets, shared descriptor offsets, same-file non-overlapping ranges, and rejection of overlapping ranges. Test filesystems with and without acceleration, same-filesystem and cross-filesystem pairs, network mounts if supported, sparse files, destination files larger than the copy range, disk-full behavior, interruption, and kernel versions at the deployment floor.
Compare destination bytes against a stable source and check the expected final length. Inspect hole allocation separately when preserving sparse layout. Record which path was used, bytes transferred, elapsed time, fallback reason, and filesystem type. Avoid logging sensitive paths or content in broadly accessible telemetry.
copy_file_range() is valuable because it gives the kernel and filesystem room to optimize file-to-file copies. Its safe use still depends on an explicit byte-range contract, correct offset accounting, a carefully selected fallback, and separate handling for metadata, sparse layout, integrity, and durability.
Related:
- Understanding the Linux Page Cache and How Writeback Actually Works
- The Linux Virtual File System: One Interface, Many Filesystems
Sources: