Linux posix_fadvise: Page-Cache Hints for Real I/O Workloads
Apply Linux posix_fadvise hints carefully for streaming, random reads, and cache release while accounting for advisory semantics and return codes.
Linux posix_fadvise() lets an application describe an expected file-access pattern so the kernel can adjust read-ahead or page-cache behavior. The advice is nonbinding: it expresses intent, not a command that the kernel must satisfy. Used thoughtfully, it can reduce wasted I/O for a sequential scan or prepare a small region before a latency-sensitive read. Used as a performance superstition, it can evict useful cache or make a workload slower.
The most important operational rule is to measure the real access pattern under realistic memory pressure. File caching is shared system state, and a hint that helps one process can harm another process using the same data. posix_fadvise() does not create a private cache or guarantee data is resident, absent, or persisted.
Scope advice to a descriptor and byte range
posix_fadvise(fd, offset, size, advice) associates an access-pattern hint with a region of the file referred to by fd. A size of zero means from the supplied offset through the current end of file. Validate that offsets are nonnegative and representable before calling the function, especially when they are computed from untrusted input or a 32-bit ABI.
#define _POSIX_C_SOURCE 200112L
#include <errno.h>
#include <fcntl.h>
#include <sys/types.h>
int advise_streaming_range(int fd, off_t offset, off_t length) {
int error = posix_fadvise(fd, offset, length, POSIX_FADV_SEQUENTIAL);
if (error != 0) {
errno = error; /* The function returns the error number directly. */
return -1;
}
return 0;
}
Unlike many system calls, posix_fadvise() returns zero on success and a positive error number on failure; it does not return -1 and set errno as its primary error convention. Wrappers that check only errno can silently ignore a failed hint. The file descriptor must refer to a seekable file; a pipe or FIFO returns ESPIPE on current Linux.
The function does not change the descriptor’s current file offset. It also does not change file contents or establish synchronization with another process. Advice may be ignored or partially applied depending on the kernel, memory pressure, filesystem, and specific hint.
Match the hint to the workload
POSIX_FADV_NORMAL states that the application has no special pattern to report. On Linux, it restores the default readahead window for the backing device. POSIX_FADV_SEQUENTIAL declares that lower offsets will be read before higher ones; Linux increases the readahead window. POSIX_FADV_RANDOM indicates random access and disables file readahead on Linux.
The names can suggest that each hint affects only the requested byte range, but Linux documents that the sequential and random readahead changes affect the entire file. The same documentation notes that other open file handles to the same file are unaffected. Coordinate through an application policy if several readers have conflicting access patterns. A database-style random reader and a streaming backup process should not casually share assumptions about one file’s cache behavior.
POSIX_FADV_WILLNEED initiates a nonblocking read of the specified region into the page cache. The kernel may reduce the amount of data read according to virtual-memory load; a request does not mean all bytes are resident by the time the function returns. It is a prefetch hint, not a completion event. If the application requires confirmed data availability, it still needs to perform and handle the actual read.
POSIX_FADV_DONTNEED asks the kernel to discard cached pages associated with a region. It is best effort and partial pages are ignored, so offsets and lengths should be page-aligned when the caller wants the full range considered. Dirty pages are not guaranteed to be written back and freed; synchronize them with fdatasync() or fsync() first if the workload specifically needs dirty cache to become eligible for release.
POSIX_FADV_NOREUSE describes one-pass access. Its Linux behavior changed over time: it was effectively a no-op for many kernel versions, and since Linux 6.3 it allows the page-replacement algorithm to ignore access to page-cache pages marked by this flag. Support and impact should be checked against the deployment kernel. Do not cite a modern NOREUSE effect as though it applied to every older Linux installation.
Do not confuse hints with direct I/O or durability
posix_fadvise() does not switch a file to O_DIRECT, bypass the page cache, or force data to storage. A sequential hint can cause more read-ahead while data still travels through the cache. A DONTNEED request can attempt eviction but does not prove that the underlying storage blocks are free of copies or that a security erase occurred.
Likewise, advising that data will be read soon is not equivalent to reading it, locking it in RAM, or reserving I/O bandwidth. The kernel can reclaim pages under pressure. If a request is a hard latency requirement, model the queue depth, filesystem, device, CPU scheduling, and memory pressure rather than depending on a soft cache hint.
posix_fadvise() is also distinct from posix_madvise() and madvise(). Those interfaces describe memory mappings, while posix_fadvise() describes file access through a descriptor. readahead() is another Linux-specific way to ask for read-ahead, while mincore() can give a limited view of resident pages in a mapping. These APIs have different scope and guarantees; choose based on the object the application is actually controlling.
Use DONTNEED carefully in shared systems
A streaming process may read a large dataset once and then mark completed ranges as unlikely to be reused, reducing the chance that its scan displaces frequently used cache. That can help a dedicated backup or batch job. But if another process is about to read the same pages, discarding them can force additional storage I/O and increase latency.
Avoid issuing DONTNEED after every small read without measuring page alignment and readahead effects. Unaligned requests may leave edge pages cached, and repeated cache churn can defeat useful readahead. Use a bounded window, issue advice after data has truly been consumed, and track the cache behavior alongside throughput and latency.
The global /proc/sys/vm/drop_caches interface is not a safe per-process tuning tool. It affects broader system caches and should not be used in production to make benchmarks appear cold. To evaluate a hint, run controlled tests in an isolated environment and document whether the result was warm-cache, cold-cache, or memory-pressure behavior.
Handle concurrency and file mutation
Advice is not a lock. Another thread or process may change the file, truncate it, rename it, or read the same region under another pattern while the hint is in effect. Keep correctness independent of cache residency and synchronize application-level content versions separately.
If an application reads a changing file as a snapshot, it needs a versioning or locking design; POSIX_FADV_SEQUENTIAL does not make the read consistent. For a streaming copy, use explicit offsets or independent open file descriptions when shared offsets would create races, and verify the byte count and checksum under the intended source-stability policy.
Network filesystems and unusual storage stacks can have their own cache and coherency layers. A client-side page-cache hint does not invalidate server caches or tell a remote writer to flush. For remote content, follow the filesystem’s consistency protocol and use polling or application-level notifications where changes must be detected.
Measure whether advice helped
Benchmark representative working sets rather than tiny synthetic files that fit entirely in memory. Vary file size, access stride, concurrency, memory pressure, storage type, and repeated versus one-pass scans. Compare no advice, sequential, random, will-need, and carefully aligned dont-need modes where applicable.
Record throughput, tail latency, cache hit behavior, read-ahead volume, I/O queue depth, memory reclaim, and the effect on competing processes. Verify the advice return code and kernel version. A faster isolated benchmark can still be a regression on a shared host if it increases cache misses for latency-sensitive services.
Test failure paths too: invalid advice, non-seekable descriptors, resource pressure, large ranges, concurrent file replacement, dirty data before DONTNEED, and a kernel where the desired behavior is unavailable or different. The result should be a measured workload-specific decision, not a blanket claim that advice always makes file I/O faster.
posix_fadvise() is a small interface with system-wide consequences. Use it to communicate a credible access pattern, check its direct error return, account for Linux-specific behavior and kernel-version changes, and keep correctness independent of the kernel honoring the hint.
Related:
- Understanding the Linux Page Cache and How Writeback Actually Works
- Linux fsync and Atomic File Replacement: Visibility Is Not Durability
Sources: