Linux renameat2: No-Replace Publication and Atomic Namespace Changes
Use renameat2 flags for no-overwrite publication and atomic exchange while handling filesystem support, directory descriptors, and crash durability.
Renaming a file changes a directory entry, not the file’s contents. Linux renameat2() extends the familiar rename operation with flags that express stronger namespace policies: do not replace an existing target, atomically exchange two names, or create an overlay-filesystem whiteout while moving an entry.
These flags are useful for package managers, deployment tools, caches, and transaction-like file updates. They do not provide a general compare-and-swap over file contents, work across different mounted filesystems, or guarantee that a successful namespace update has reached stable storage. Callers still need correct directory descriptors, a recovery policy, and explicit durability steps.
Use directory descriptors to bound name resolution
renameat2() takes an old directory descriptor and relative path, a new directory descriptor and relative path, and a flags value. Relative names are resolved from the corresponding directory descriptors rather than from a mutable process working directory. This is useful in a privileged helper that should operate inside a pre-opened directory.
#define _GNU_SOURCE
#include <fcntl.h>
#include <stdio.h>
int publish_if_absent(int dirfd, const char *temporary_name,
const char *final_name) {
return renameat2(dirfd, temporary_name, dirfd, final_name,
RENAME_NOREPLACE);
}
This helper returns zero when the rename succeeds and minus one with errno set when it fails. RENAME_NOREPLACE returns EEXIST if the destination already exists instead of overwriting it. The flag is supported only when the running kernel and the filesystem implement it; a kernel or filesystem can reject the flag with EINVAL.
Use a directory descriptor acquired from a trusted parent and validate each name as a single intended path component when the application expects that. Directory descriptors reduce reliance on the process working directory, but they do not automatically prevent .., symlink, or mount traversal in every pathname design. If a path crosses an attacker-controlled tree, combine the rename operation with an explicit path-resolution policy such as openat2() on the descriptors or a directory structure that cannot be replaced by an untrusted user.
Avoid check-then-rename races
A fragile no-overwrite implementation first checks that a destination does not exist and then calls ordinary rename(). Another process can create the destination between those operations, and ordinary rename may replace it. RENAME_NOREPLACE combines the existence condition with the namespace operation so the kernel decides the winner atomically on supporting filesystems.
This is appropriate for create-once artifact publication, unique cache entries, lock-file names, and installer generations. It does not prove that the temporary file is correct or that the caller is authorized to publish every object in the directory. A successful operation means that the name was not present at the atomic decision point and that the rename passed normal directory permission and filesystem checks.
If the filesystem does not support RENAME_NOREPLACE, do not silently emulate it with access() followed by rename(). Use a different atomic primitive that matches the filesystem and application contract, fail closed, or explicitly downgrade to a race-prone mode with clear documentation. A hard-link-based publication scheme may be viable for some regular files but has different semantics and limitations; it is not a universal fallback.
Exchange two existing names atomically
RENAME_EXCHANGE atomically swaps the objects named by the source and destination paths. Both names must exist. The objects can have different types, such as a non-empty directory and a symlink, subject to filesystem support. This can support staged configuration generations, directory-tree rollouts, and two-slot caches where readers should see one complete generation or the other.
An exchange does not merge the objects or preserve a higher-level invariant by itself. Open descriptors continue to refer to the objects they already opened. Future path lookups see the exchanged names. Readers that open several files independently can still observe different generations if an exchange occurs between opens; use a single manifest, generation directory descriptor, or application-level snapshot protocol for multi-file consistency.
Do not use RENAME_EXCHANGE as if it were a lock. Two writers can exchange repeatedly or race with other directory operations. If the application requires a particular expected generation, record and verify a generation identity under a lock or other coordination protocol before and after the exchange.
Understand whiteout as an overlay-filesystem operation
RENAME_WHITEOUT is intended for overlay or union filesystem implementations. It moves the upper-layer object and creates a whiteout at the source so a matching lower-layer name stays hidden. The operation is atomic for the supported overlay semantics, but a whiteout outside such a context appears as a special character device-like object rather than an ordinary missing file.
Creating a whiteout requires the relevant privilege, including CAP_MKNOD, and filesystem support. It cannot be combined with RENAME_EXCHANGE. Most application code should not set this flag. It belongs in filesystem and container-layer implementations that understand how lower and upper trees are composed.
Distinguish atomic visibility from crash durability
A rename within one mounted filesystem gives atomic namespace visibility: observers do not see a partially copied directory entry. If replacing a file safely, a common sequence writes and synchronizes a temporary file, renames it into place, and synchronizes the containing directory. The file synchronization persists content and associated metadata; the directory synchronization is needed to persist the name update on filesystems where directory entries require it.
renameat2() returning success does not mean data is durable after power loss. If the caller needs that guarantee, use the appropriate fsync() steps and test the actual filesystem and storage stack. For more complex changes involving two directories, source and destination directory synchronization may both matter. For temporary files, clean up only names the application created and can identify safely.
Network filesystems add uncertainty. On NFS, a rename may have been performed by the server even when the client later observes an error after a server restart and retransmission. After an ambiguous failure, reconcile the directory state and an application-level transaction identifier rather than assuming that failure means the operation did not happen.
Respect mount boundaries and failure classes
The source and destination must be on the same mounted filesystem. A rename across mount points fails with EXDEV, even if the same underlying filesystem is mounted at both locations. Copying bytes and deleting the old name is not an atomic replacement. If cross-filesystem movement is required, define a staged copy, checksum, synchronization, publication, and cleanup protocol.
Separate EEXIST from ENOENT, EXDEV, EACCES, EROFS, EINVAL, quota errors, and I/O errors. These failures have different recovery meaning. EINVAL can indicate conflicting or unsupported flags, while EXDEV identifies a mount boundary. Blindly retrying with flags set to zero can silently change “must not overwrite” into “overwrite allowed.”
Feature support has several layers: libc wrapper availability, build headers defining flags, running kernel recognizing the operation, and filesystem supporting the requested semantics. Test the actual operation on the target filesystem and record unsupported modes. Do not assume that a test on a developer’s ext4 filesystem predicts behavior on a network mount, overlay, or older embedded kernel.
Verify the namespace contract under contention
Run concurrent publishers that all attempt RENAME_NOREPLACE for one destination and prove exactly one succeeds. For exchange, run readers while swapping two known generations and confirm each opened object is internally complete. Test missing source and destination, permission failures, read-only filesystems, mount boundaries, unsupported flags, and filesystem crash-recovery behavior.
Log the source and destination generation identifiers, directory identity, flag mode, and outcome. Avoid logging secrets in path names. If rename succeeds but a later directory sync fails, mark the update visible-but-durability-uncertain and reconcile on restart; do not report a clean rollback.
renameat2() makes several namespace policies atomic where the kernel and filesystem support them. Choose flags to match the exact publication invariant, operate relative to trusted directory descriptors, and keep visibility, durability, authorization, and multi-file consistency as separate parts of the design.
Related:
- Linux fsync and Atomic File Replacement: Visibility Is Not Durability
- Linux openat2: Constraining Path Resolution Against Symlink and Mount Escapes
Sources: