Skip to content
LinuxHow-To Published Updated 7 min readViews unavailable

Linux fsync and Atomic File Replacement: Visibility Is Not Durability

Build a crash-conscious Linux file replacement sequence with fsync, same-filesystem rename, directory synchronization, and explicit recovery rules.

Writing a new file and renaming it over an old pathname is a common way to avoid exposing a partially written configuration, index, or state snapshot. The rename can provide atomic namespace visibility when the source and destination are on the same filesystem, but atomic visibility is not the same as crash durability. A process that needs to survive power loss must order file synchronization and directory synchronization deliberately.

The exact guarantee depends on the filesystem, storage device, mount configuration, and whether the storage stack correctly honors flush requests. fsync() is a necessary part of many durable update protocols, but the syscall name alone is not a complete end-to-end guarantee.

Separate data persistence from name persistence

A file has data and metadata associated with its inode, and a directory has entries that map names to objects. fsync(file_fd) flushes modified file data and associated metadata for that open file. It does not necessarily make the directory entry that names the file durable. A newly created or renamed name may require an explicit fsync() on an open descriptor for the containing directory.

This distinction creates two different milestones. First, the new file’s contents and required metadata are synchronized. Second, the directory update that publishes the final name is synchronized. A crash between those steps can leave a different namespace outcome than the application intended, even if the write calls returned success.

fdatasync() can avoid synchronizing metadata that is not needed for subsequent data retrieval. It still needs to persist metadata such as file size when that size is necessary to read the written bytes. Use fsync() when the application requires the broader file metadata update, and use fdatasync() only when its narrower contract matches the file format’s durability needs.

Write to a temporary file, then publish it

A robust replacement design creates a temporary file in the destination directory, writes the complete new content, synchronizes it, renames it over the final path, and synchronizes the directory. Keeping the temporary file in the same directory helps ensure the rename stays on one filesystem and makes the final namespace operation easier to reason about.

#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <limits.h>
#include <stddef.h>
#include <stdio.h>
#include <sys/types.h>
#include <unistd.h>

static int write_all(int fd, const void *buffer, size_t length) {
    const unsigned char *p = buffer;
    while (length > 0) {
        size_t request = length > (size_t)SSIZE_MAX
                       ? (size_t)SSIZE_MAX : length;
        ssize_t n = write(fd, p, request);
        if (n > 0) {
            p += (size_t)n;
            length -= (size_t)n;
            continue;
        }
        if (n == -1 && errno == EINTR)
            continue;
        if (n == 0)
            errno = EIO;
        return -1;
    }
    return 0;
}

int publish_snapshot(int dirfd, const char *temporary_name,
                     const char *final_name, const void *data, size_t length) {
    int fd = openat(dirfd, temporary_name,
                    O_WRONLY | O_CREAT | O_EXCL | O_CLOEXEC | O_NOFOLLOW, 0600);
    if (fd == -1)
        return -1;

    if (write_all(fd, data, length) == -1 || fsync(fd) == -1) {
        int saved_errno = errno;
        (void)close(fd);
        (void)unlinkat(dirfd, temporary_name, 0);
        errno = saved_errno;
        return -1;
    }
    if (close(fd) == -1) {
        int saved_errno = errno;
        (void)unlinkat(dirfd, temporary_name, 0);
        errno = saved_errno;
        return -1;
    }

    if (renameat(dirfd, temporary_name, dirfd, final_name) == -1) {
        int saved_errno = errno;
        (void)unlinkat(dirfd, temporary_name, 0);
        errno = saved_errno;
        return -1;
    }
    /* If this fails, publication happened but crash durability is uncertain. */
    return fsync(dirfd);
}

The function assumes the directory descriptor refers to a trusted directory and that the filesystem and application permit replacing the destination. Its return value alone does not distinguish every failure stage: a failure from the final directory fsync() means rename already happened and crash durability is uncertain. A production API should return a stage or transaction identifier so the caller can reconcile that state. Once rename succeeds, unlinking the old temporary name does not undo publication because that name no longer refers to the new file.

If the temporary and final names are in different directories on the same filesystem, the rename updates both directories. A crash-safe protocol may need to synchronize both directory descriptors after the operation. If they are on different mounted filesystems, rename fails with EXDEV; copying bytes and deleting the source is not an atomic substitute.

Understand the atomicity boundary

For a same-filesystem rename, observers see the old name binding or the new one, rather than a partially copied destination entry. Existing open descriptors to the old file continue to refer to that old inode, while future path lookups resolve according to the renamed directory entry. Readers that need a consistent view should open the file once and keep that descriptor for the entire read.

Atomic rename does not serialize multiple writers. Two processes can each write and synchronize a temporary file, then rename over the same destination; the last successful namespace operation wins. If lost updates are unacceptable, use a lock, compare-and-swap style generation file, RENAME_NOREPLACE where appropriate, or a transactional storage engine.

Atomic visibility also does not mean that directory updates have reached persistent media. The directory fsync() after rename is part of the durability protocol. Some filesystems reject directory synchronization or have different guarantees; test the actual target filesystem and document any weaker mode. Do not report “durable save complete” before the required synchronization calls have succeeded.

Handle failures at each state transition

Track the transaction stages: temporary file created, bytes written, file synchronized, descriptor closed, name replaced, directory synchronized. If a write fails, the final name should remain unchanged and the temporary file should be removed or left for safe startup cleanup. If file synchronization fails, do not publish the new version unless the application explicitly accepts a weaker guarantee.

If rename succeeds but directory synchronization fails, the new file may already be visible to running processes, but crash persistence is uncertain. The caller should report an uncertain-commit state instead of pretending the operation rolled back. Recovery logic can inspect a generation number or checksum at startup and determine which complete version exists.

fsync() can report errors from earlier writeback, including storage errors associated with the file. On Linux, a reported EIO may not identify only the most recent write or one descriptor when multiple descriptors share a file or storage device. Treat synchronization failure as serious, preserve diagnostic context, and do not retry indefinitely without a recovery policy.

If the process crashes before removing an uncommitted temporary file, a startup scan can remove only files that match a strict naming and ownership pattern. Use random or collision-resistant names and create them with exclusive creation. Never delete arbitrary files in the directory simply because their names share a weak prefix.

Keep the temporary object private until ready

Create the temporary file with restrictive permissions, then apply the exact intended mode, ownership, ACLs, extended attributes, and security labels before publication. Decide when each metadata change must be synchronized. Copying all source metadata can transfer privileges or labels that the replacement process should not retain.

For secret-bearing configuration, prevent other users from opening the temporary file before permissions are finalized. A directory with writable access by untrusted users can permit name substitution and cleanup attacks; use a trusted directory descriptor and avoid re-resolving a path through an attacker-controlled working directory. If the application needs a no-overwrite publication, Linux renameat2() with RENAME_NOREPLACE can express that policy on supporting filesystems, but filesystem support must be checked.

Memory-mapped writes, direct I/O, delayed allocation, copy-on-write filesystems, and network filesystems can add their own synchronization requirements. Do not assume that a library’s “save” call uses the exact ordering your durability contract needs. Trace or inspect its behavior when data-loss resistance is critical.

Test crash outcomes, not only return codes

Test process termination at every stage and, where possible, inject filesystem or block-device failures. Verify which version is visible after restart, whether an interrupted temp file is detectable, and whether the directory entry survives. A unit test that observes a successful return in a running kernel does not simulate a power failure or broken storage flush implementation.

Use a filesystem-specific test environment for ext4, XFS, Btrfs, or the remote filesystem in production. Include full disks, quota exhaustion, read-only remounts, EIO, rename replacement, concurrent writers, and directory fsync() behavior. A VM or loopback filesystem is useful for repeatability but may not model a storage controller’s volatile cache.

Log the transaction ID, stage reached, synchronization result, and recovery decision. Avoid logging file contents or secrets. On ambiguous failure after rename, mark the operation as uncertain and reconcile the durable state before accepting a competing update.

The safe mental model is simple: write and synchronize the new inode, atomically change the directory entry, then synchronize the directory. Every filesystem and application adds details, but none of those details make visibility alone equivalent to durability.

Related:

Sources:

Comments