Linux fallocate: Reserve, Punch, Zero, and Reshape File Space
Use fallocate deliberately for space reservation, sparse-file hole punching, zero ranges, and range changes while accounting for filesystem support.
Linux fallocate() lets software ask a filesystem to manipulate the space associated with a regular file. The basic operation can allocate backing space for a range before the application writes it. Flagged forms can keep the visible file size unchanged, create holes, make a range read as zeros, or insert and remove ranges that shift later bytes.
These operations are not interchangeable. A reservation is not a durability guarantee, a logical zero range is not secure erasure, and punching a hole is not the same as truncating a file. The correct mode depends on how the application defines file size, sparse layout, concurrent access, and failure recovery.
Reserve space before a write-heavy phase
The ordinary fallocate(fd, 0, offset, length) operation asks the filesystem to allocate space for the specified byte range. This can move allocation failure earlier, before the application commits to a transaction or starts a long write. On success, subsequent writes within the allocated range are protected from failing solely because the filesystem runs out of space, although they can still fail for other reasons such as I/O errors or policy restrictions.
The descriptor must refer to a file opened for writing, and offset + length must be representable and within the filesystem’s file-size limits. The range need not fit the file’s current logical size: ordinary preallocation may extend it, while FALLOC_FL_KEEP_SIZE can reserve beyond the end without changing the reported size. These are file-space operations, not a way to reserve storage for a pipe, socket, or arbitrary device.
#define _GNU_SOURCE
#include <fcntl.h>
#include <unistd.h>
int reserve_for_append(int fd, off_t start, off_t length) {
if (start < 0 || length <= 0)
return -1;
return fallocate(fd, 0, start, length);
}
Validate that start + length is representable and within application and filesystem limits before calling the syscall. A successful ordinary allocation may extend the file size; reading unwritten portions returns zeros. Allocation is performed in filesystem block units, so the filesystem may reserve more than the exact byte length requested.
For an append-oriented format that wants to reserve blocks beyond the current end without changing the logical size yet, use FALLOC_FL_KEEP_SIZE where supported. This separates physical reservation from the file’s visible length. The application must later write and update size according to its format, and it should not expose bytes beyond the committed logical boundary merely because blocks were reserved.
Space reservation is a capacity planning tool, not a promise that a storage device will survive hardware failure or that metadata has reached stable storage. If the filesystem is copy-on-write, thin-provisioned, remote, or subject to quotas, understand its allocation model and the point at which capacity is actually guaranteed. The syscall can fail with ENOSPC, EDQUOT, EOPNOTSUPP, or other errors; handle these before publishing a success state.
Punch a hole without changing file length
FALLOC_FL_PUNCH_HOLE deallocates storage in a byte range while retaining the file’s length. It must be combined with FALLOC_FL_KEEP_SIZE. Reads from the punched range return zeroes. Full filesystem blocks can be removed from allocation, while partial boundary blocks are zeroed so bytes outside the requested range remain intact.
Hole punching is useful for sparse images, append-only logs with expired segments, and storage formats that can discard unused regions. It does not shift later bytes. Applications with an index must continue to interpret offsets exactly as before; the content in the punched region has become zero-filled from the file’s logical perspective.
Not every filesystem supports punching holes, and supported behavior can differ by filesystem and kernel. Treat EOPNOTSUPP or EINVAL as a capability or request-geometry result that deserves explicit policy. Do not silently rewrite a sparse file to a fully allocated file as a fallback if saving space is a requirement. The fallback might preserve bytes but violate the storage contract or exhaust the volume.
Hole punching is also not a secure erase primitive. Filesystem snapshots, reflinks, journals, backups, block-layer remapping, flash translation layers, and remote replicas may retain previous data or copies. It makes the live file read as zeros and can release logical allocation; it does not prove every physical copy has been destroyed.
Zero a range and preserve its logical contents
FALLOC_FL_ZERO_RANGE makes reads in the selected range return zeroes while generally allocating blocks for the range. Filesystems may represent that state as unwritten extents and avoid writing all zero bytes to the storage device. This can be far faster than issuing a write of a large zero-filled userspace buffer.
The semantic difference from punching is allocation intent: zero-range operations normally retain or create allocated space for the range, while punching aims to deallocate it. This can be useful when a file format requires a zero-filled region and future writes should not encounter a space shortage solely because those blocks were never reserved. With FALLOC_FL_KEEP_SIZE, a range can be prepared beyond the current file length without extending the reported size.
Do not infer physical zeroing from logical zero reads. Filesystem extent metadata may satisfy the reads, and snapshots or copy-on-write layers can retain prior versions. If a cryptographic or regulatory erasure guarantee is required, define it against the storage stack, key management, snapshots, and backups instead of treating ZERO_RANGE or a zero-filled write as erasure.
Insert or collapse a range only when offsets may move
FALLOC_FL_INSERT_RANGE inserts a hole inside an existing file and shifts later bytes toward larger offsets. The file grows by the inserted length. FALLOC_FL_COLLAPSE_RANGE removes bytes and shifts subsequent data toward smaller offsets, shrinking the file. These are structural edits to the file’s logical byte sequence, not just allocation changes.
The operations have filesystem and alignment restrictions. Some require offset and length to align to filesystem block boundaries; unsupported filesystems reject them. They cannot be combined with arbitrary other flags. Range collapse cannot extend through or beyond end-of-file; use ftruncate() for ordinary truncation at the end. Range insertion at or beyond end-of-file is likewise not a substitute for extending the file.
Shifting bytes invalidates external offsets, indexes, memory maps, checksums, and concurrent readers that assume a stable layout. Lock the file or coordinate through a higher-level transaction before changing ranges. After a successful insertion or collapse, update metadata that stores offsets, then verify the resulting file size and content. If the process can crash between the filesystem operation and index update, recovery needs a journal or a rebuildable index.
Handle filesystem variance without hiding errors
The fallocate() interface is Linux-specific, and flags have different kernel and filesystem support histories. Build against the oldest supported headers or guard newer constants at compile time, then detect unsupported behavior at runtime. A filesystem may support ordinary preallocation but not hole punching, zeroing, collapse, or insertion.
Keep capability fallback separate from storage failure. If a reserve call returns ENOSPC, silently switching to a buffered write does not make the capacity constraint disappear. If hole punching is unsupported, a zero-write fallback can consume more space and does not produce the same sparse layout. Record which operation was attempted and which fallback policy was selected.
Treat concurrent access as a correctness concern. The syscall changes file allocation and potentially length while other file descriptors may be writing or reading the same inode. A mutex in one process does not coordinate with other processes unless they share a protocol. Use file locks or an application-level transaction when operations must be serialized, and consider leases or immutable versioning for remote systems where local locking is insufficient.
Inspect the result using the right measurements
stat() reports logical file size and allocated blocks, but the unit and representation of allocation information need careful interpretation. Filesystems can compress, deduplicate, share, delay, or otherwise transform allocation. Sparse layout can be queried with SEEK_DATA and SEEK_HOLE where supported, but those results are filesystem views rather than a forensic statement about every physical block.
Test on each deployment filesystem and include normal files, sparse files, reflinked files, quota limits, full volumes, alignment boundaries, concurrent readers, and injected I/O errors. Verify logical bytes with reads or hashes, visible length with metadata, and allocation using the filesystem-appropriate inspection tools. Measure after synchronization only if the requirement is about persisted state, and test recovery across a crash when the operation participates in a transaction.
Use ordinary preallocation to surface capacity problems before writes, punch holes to release logical ranges where supported, zero ranges to establish logical zeros efficiently, and insert/collapse operations only when shifting all later offsets is intended. fallocate() is powerful precisely because these effects are distinct; selecting the wrong flag can preserve bytes but break the file’s storage or indexing contract.
Related:
- Fixing ‘Disk Full’ on Linux When df Shows Space Available
- Linux copy_file_range: Efficient File Copies with Explicit Boundaries
Sources: