Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux cachestat: Inspect File Page-Cache Residency by Range

Use Linux cachestat to inspect cached, dirty, writeback, and evicted pages for a file range, then interpret the snapshot without confusing it with I/O telemetry.

When a file-backed workload slows down, operators often need to know whether the relevant pages are resident in the Linux page cache, being written back, or recently evicted under memory pressure. System-wide counters such as /proc/meminfo do not identify a file range, while a mapped-page probe such as mincore() requires an address mapping and does not return dirty or writeback counts.

Linux cachestat() provides a file-descriptor-based snapshot for a byte range. It reports counts for cached, dirty, writeback, evicted, and recently evicted pages. The syscall was added in Linux 6.5. It is useful for targeted diagnosis and cache-aware tooling, but it is not a disk-I/O counter, a per-process resident-set report, or a guarantee that a subsequent read will hit memory.

What the five counters mean

The call takes an open file descriptor, a range described by byte offset and length, an output structure, and a flags value that must currently be zero. A positive length queries the requested interval; a zero length means from the starting offset to the end of the file. On success, the kernel fills an aggregate result rather than returning one bit per page.

The fields distinguish several states that are operationally different:

Field Diagnostic meaning
nr_cache Pages in the queried range that are currently represented in page cache.
nr_dirty Cached pages that have modifications not yet written back.
nr_writeback Pages currently marked for writeback.
nr_evicted Evicted-page metadata found for the queried range.
nr_recently_evicted The subset classified as recently evicted, a signal that re-entry could indicate active use under memory pressure.

These are range-level observations. They do not say which process caused a page to enter the cache, which process will use it next, or how many physical device reads occurred. Multiple processes can share the same inode’s page cache, and the status can change while the call runs or immediately afterward. Treat the values as a timestamped sample, not a transactionally consistent snapshot.

Query a specific range through the syscall API

On systems whose Linux UAPI headers expose the structures and syscall number, a small diagnostic helper can invoke the syscall through syscall(2). This avoids assuming that every deployed C library version provides a direct cachestat() wrapper. Compile against the target system’s Linux headers, use its architecture-provided SYS_cachestat definition, and keep a fallback for ENOSYS or a policy error.

#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <linux/mman.h>
#include <stdio.h>
#include <sys/syscall.h>
#include <unistd.h>

int main(int argc, char **argv) {
    if (argc != 2) {
        fprintf(stderr, "usage: %s FILE\n", argv[0]);
        return 2;
    }

    int fd = open(argv[1], O_RDONLY | O_CLOEXEC);
    if (fd == -1) {
        perror("open");
        return 1;
    }

    struct cachestat_range range = { .off = 0, .len = 0 };
    struct cachestat stats = {0};
    long rc = syscall(SYS_cachestat, fd, &range, &stats, 0);
    if (rc == -1) {
        int saved_errno = errno;
        perror("cachestat");
        close(fd);
        return saved_errno == ENOSYS ? 3 : 1;
    }

    printf("cached=%llu dirty=%llu writeback=%llu evicted=%llu recently_evicted=%llu\n",
           (unsigned long long) stats.nr_cache,
           (unsigned long long) stats.nr_dirty,
           (unsigned long long) stats.nr_writeback,
           (unsigned long long) stats.nr_evicted,
           (unsigned long long) stats.nr_recently_evicted);
    close(fd);
    return 0;
}

The example requests the whole file from offset zero. For a bounded range, set off and len using checked 64-bit arithmetic; reject negative values and guard off + len against overflow before issuing the call. Keep the file descriptor open for the duration of the query so the inode being measured is stable. The output counts pages, not bytes; do not multiply by a hard-coded 4096 on systems where the kernel page size differs. If a byte estimate is useful, obtain the runtime page size and label the result as approximate.

SYS_cachestat and the UAPI structure definitions are architecture and header dependent. A missing macro is a build-environment problem, not proof that the running kernel lacks the syscall. Conversely, successful compilation does not prove that the runtime kernel implements it. Linux 6.5 is the upstream introduction point; older kernels return ENOSYS, and seccomp or another syscall policy may reject the call separately. Do not copy one architecture’s syscall number into a portable program.

Interpret samples alongside workload evidence

A high nr_cache means many pages in that file range were cached when queried. It does not establish that the application’s next access will be served from cache: pages can be reclaimed or invalidated, the workload can move to a different file range, and reads can be performed with flags or I/O paths that bypass ordinary buffered page-cache behavior. Compare the sampled range with the application’s actual offsets and access pattern.

Dirty and writeback counts help separate cache residency from persistence pressure. A file can be largely cached while writes are still waiting for writeback. A rising writeback count may coincide with device congestion, but it is not by itself proof that the storage device is saturated. Correlate samples with process I/O, filesystem state, device latency, and the application’s write and sync behavior. Avoid calling sync or dropping caches just to make the counters look clean; those actions alter the workload and can hide the condition being diagnosed.

Evicted and recently evicted values provide context about reclaim. They should not be added to nr_cache and described as total file reads, nor converted into a cache-hit ratio without a measurement model that tracks accesses over the same interval. A repeated sequence of snapshots can show whether the range’s observed cache state is changing, but a monitoring agent should record timestamps, range boundaries, file identity, kernel version, and errors so results remain comparable.

For a minimal operational check, run the diagnostic before and after the workload phase of interest and keep the sampled interval fixed:

file=/var/lib/example/index.db
stat --printf='device=%D inode=%i size=%s\n' "$file"
./cachestat-range "$file"

On a Linux build host, strace -e trace=cachestat can confirm that the program reached the syscall and show the returned status. Do not benchmark through strace, and do not infer page-cache state from an absent trace if the tool does not recognize the syscall name. Record uname -r and the relevant libc and header package versions when diagnosing an environment mismatch.

Know the boundaries before automating it

cachestat() is file- and range-oriented. It is not a direct view of a process’s resident set, and its aggregated results do not identify individual page offsets. Use /proc/PID/smaps or mapping-level tools for process mapping questions, device metrics for physical I/O, and application instrumentation for request-level cache outcomes. mincore() remains useful when the question is page residency in a particular mapping; cachestat() is useful when the question is aggregate cache state for a file range.

The interface currently does not support hugetlbfs files. Treat EOPNOTSUPP as a capability limitation, not as an empty-cache result. Check the target kernel and filesystem behavior in the environments you support, including containers where seccomp policy may differ from the host. Since the counters are sampled and can become outdated before user space receives them, alert on sustained patterns or corroborated symptoms rather than a single threshold crossing.

Use cachestat() to answer a narrow question: what page-cache state did the kernel report for this file range at this moment? Combine that answer with workload and device telemetry, and avoid presenting it as an exact history of I/O. That discipline makes the syscall a useful observability signal instead of a misleading performance metric.

Related:

Sources:

Comments