Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux statx: Capability-Aware File Metadata and Mount Identity

Request Linux file metadata with statx, check returned mask bits, and handle optional creation time, mount IDs, and direct-I/O alignment safely.

Linux statx() is an extensible interface for querying file metadata. It can return familiar values such as mode, ownership, size, inode number, and timestamps, plus optional information including creation time, mount identity, direct-I/O alignment, subvolume identifiers, and atomic-write limits on supported kernel and filesystem combinations.

The interface is deliberately capability-aware: asking for a field does not guarantee that the kernel or backing filesystem can provide it. Correct callers inspect the returned mask before using optional fields. Treating an unreturned field as zero can turn “not supported” into a false claim about the file.

Request only the fields the caller needs

statx() takes a directory descriptor, a pathname, lookup flags, a request mask, and an output structure. The *at()-style directory descriptor makes it possible to resolve a relative path from an already-open directory instead of depending on the process working directory. Flags such as AT_SYMLINK_NOFOLLOW let a caller choose whether it is describing the symlink itself or its target.

#define _GNU_SOURCE
#include <fcntl.h>
#include <stdio.h>
#include <sys/stat.h>
#include <unistd.h>

int inspect_path(int dirfd, const char *path) {
    struct statx info = {0};
    unsigned int wanted = STATX_BASIC_STATS | STATX_BTIME;

    if (statx(dirfd, path, AT_SYMLINK_NOFOLLOW, wanted, &info) == -1)
        return -1;

    if (info.stx_mask & STATX_TYPE)
        printf("mode type bits: %o\n", info.stx_mode & S_IFMT);
    if (info.stx_mask & STATX_SIZE)
        printf("logical size: %llu\n",
               (unsigned long long)info.stx_size);
    if (info.stx_mask & STATX_BTIME)
        printf("creation-time metadata is available\n");
    else
        printf("creation time is unavailable\n");
    return 0;
}

The example requests basic statistics and birth time, but it still checks each returned mask bit before using the field. In a real program, avoid printf() as a production logging strategy for untrusted paths, and format 64-bit fields using portable integer-format macros when needed. The key rule is that info.stx_mask is the authority for which fields were filled in.

Use a narrow request mask rather than setting every bit. New mask bits can be defined in later headers, and the kernel may treat unknown requested bits as extension space. STATX_ALL is deprecated; select the exact values the caller consumes. The returned mask is not required to equal the requested mask: the kernel may provide additional fields or omit requested fields that the filesystem cannot supply.

Distinguish unsupported metadata from a zero value

Creation time (STATX_BTIME) is not universally supported. An absent mask bit means the caller does not have a valid birth time for that object; it does not mean the object was created at the Unix epoch. The same principle applies to newer fields such as direct-I/O alignment and write-atomic limits.

This is important for backup tools, inventory databases, and cache keys. If a field is optional, represent “not available” explicitly in the data model. Do not serialize an uninitialized field, a zero-filled placeholder, or a stale value from an earlier query as though the current filesystem confirmed it.

A metadata query is also a point-in-time observation, not a lock. The file can be renamed, replaced, or changed after the call. If correctness depends on operating on the exact object whose metadata was inspected, open it and query the descriptor with an appropriate AT_EMPTY_PATH pattern where supported, or verify identity again through the descriptor you will use. A separate path-based check followed by an open can create a time-of-check/time-of-use race.

Use mount IDs for mount-aware diagnostics

STATX_MNT_ID can identify the mount containing the file, which is useful when a pathname crosses bind mounts, container mount namespaces, or multiple views of a filesystem. This can explain why a path resolves to an unexpected storage layer even when device numbers alone are not enough to distinguish mount instances.

Linux 6.8 added STATX_MNT_ID_UNIQUE, which requests a unique mount ID that is not reused during the system’s lifetime. Ordinary mount IDs and unique mount IDs serve different purposes; choose the one that matches the correlation lifetime. A mount ID is not a globally persistent identifier across reboots and should not be stored as a durable identity for a file or filesystem.

If the requested mount-ID bit is absent from stx_mask, do not trust the corresponding field. Newer headers may define a mask that an older runtime kernel does not support, and the backing filesystem or mount context can affect returned metadata. Feature-test the mask and interpret the result rather than assuming the build host and production host are identical.

Mount identity is also not authorization. Knowing that a file belongs to a mount does not prove the caller can read it, that the mount is trusted, or that a path cannot change after inspection. Use mount IDs for diagnostics and policy inputs only in a larger design that validates descriptor ownership and path resolution.

Discover direct-I/O requirements instead of guessing

STATX_DIOALIGN asks for memory-buffer and file-offset alignment requirements for direct I/O. Support varies by kernel and filesystem. Linux 6.1 support exists for block devices and for regular files on ext4, f2fs, and XFS; other filesystems and versions may omit the fields. A zero alignment value indicates that direct I/O is not supported or that the relevant field is not available, so check both the requested result bit and the returned values.

Linux 6.14 added a separate direct-I/O read-offset alignment field through STATX_DIO_READ_ALIGN; it can refine read requirements on supported filesystems. Direct-I/O constraints are not safely inferred from page size or sector size alone. Query the file, allocate aligned buffers, align offsets and transfer lengths, and still handle runtime I/O errors.

The API has continued to grow. STATX_SUBVOL and STATX_WRITE_ATOMIC are newer requests, with filesystem-dependent support. Headers can expose constants before all production kernels support them, and support can be backported. Avoid version-string checks as a substitute for asking the actual file and checking the returned mask.

Choose synchronization behavior for the metadata query

statx() supports synchronization flags that let callers express whether they require metadata synchronized with the filesystem or can accept cached information, according to the documented AT_STATX_SYNC_AS_STAT, AT_STATX_FORCE_SYNC, and AT_STATX_DONT_SYNC semantics. This choice can matter on network filesystems or systems where freshness has cost.

Do not describe a cached answer as a strongly consistent snapshot. Conversely, forcing synchronization for every inventory walk can make a large scan unexpectedly slow or overload a remote metadata server. Define the freshness requirement per use case: a UI listing, forensic collection, security boundary, and replication controller can require different behavior.

Permissions and path resolution still apply. A failed query can mean missing search permission, an invalid directory descriptor, an inaccessible path component, or an unsupported syscall on an older kernel. If falling back to stat() or fstatat(), document which fields and lookup semantics are lost instead of quietly returning a partial structure under the same confidence level.

Build portable code around an evolving Linux structure

The structure and mask constants evolve with Linux headers. Compile against the deployment sysroot, guard newer fields and constants when supporting older headers, and keep runtime detection separate from compile-time availability. If statx() is unavailable (ENOSYS), a fallback to fstatat() can retrieve traditional metadata but not every extension. Make this limitation visible in the result.

Do not copy the entire kernel-facing structure as a persistent binary record. Its layout, padding, and newly appended fields are not a stable application file format. Convert the values your program needs into a versioned internal representation and record which fields were actually present.

For security-sensitive decisions, combine metadata with descriptor-based operations and constrained path lookup. statx() can report mode, attributes, mount information, and timestamps, but it does not prevent another process from replacing a path afterward. Open the intended object, validate the descriptor, and operate on that descriptor rather than trusting a prior pathname query.

Test optional fields on real filesystems

Test regular files, directories, symlinks with and without following, network filesystems, bind mounts, container namespaces, and filesystems that omit birth time or direct-I/O information. Test older runtime kernels with newer headers and ensure the program handles unsupported mask bits without treating fields as valid.

For direct I/O, verify alignments using both accepted and rejected buffer/offset combinations on the actual target filesystem. For mount-aware diagnostics, compare the reported mount ID with the intended mount namespace and mount information. For freshness-sensitive workflows, measure latency under both cached and forced synchronization policies.

statx() provides a richer vocabulary for file metadata, not a guarantee that every field exists or remains true after the call. Request only what is needed, inspect the returned mask, keep optional values optional, and bind critical decisions to open descriptors rather than stale path observations.

Related:

Sources:

Comments