FreeBSD fsync and fdatasync: Define What Durable Actually Means
Distinguish FreeBSD fsync, fdatasync, sync, and user-space flushes, handle I/O errors, and test durability without overstating guarantees.
“The write returned successfully” and “the application can recover this update after a crash” are not equivalent statements. Data may move through a language runtime’s buffers, the kernel’s buffer cache, filesystem metadata operations, a device queue, and a controller or drive cache. FreeBSD exposes synchronization calls at some of these boundaries, but none of them can repair an application protocol that writes records in an unsafe order or a storage device that falsely reports completion.
This article focuses on the FreeBSD interfaces fsync(2), fdatasync(2), and sync(2). It is not a guarantee matrix for every filesystem, controller, hypervisor, or drive. Read the manual for the target release and filesystem, check device documentation, and test the actual failure model. Do not infer power-loss safety from a command name or from a clean shutdown alone.
Follow the data path before choosing an API
For a stdio stream, fprintf() can update a process-local buffer and return before a system call writes those bytes. fflush() asks the C library to pass buffered output to the underlying file descriptor; it does not by itself request that the filesystem synchronize the file to the permanent storage device. A low-level write() crosses the process/kernel boundary, but a successful return is still not the same as a successful fsync().
The basic sequence for a simple append-only diagnostic log might be:
- Format a complete record and check its length.
- Write all bytes, handling short writes and interruption.
- Call fsync() on the file descriptor at the application-defined durability boundary.
- Check the call’s result and preserve errors for operations or monitoring.
The precise transaction must be designed by the application. If one logical update spans several files, a sync on one descriptor does not make the entire multi-file update atomic. If the process crashes halfway through a record, the reader still needs framing, checksums, sequence numbers, or a journal to recognize incomplete state.
What FreeBSD documents for fsync and fdatasync
The FreeBSD 15.1 fsync(2) manual says fsync() causes modified data and attributes of the referenced file to be moved to a permanent storage device, normally resulting in modified in-core buffers for that file being written. It returns zero on success or -1 with errno set on failure. A successful result is useful evidence that the system’s synchronization path reported completion; it is not a proof that every component below the kernel has truthful power-loss behavior.
FreeBSD’s documented fdatasync() distinction is especially important. Its manual says the call moves modified data to permanent storage but does not guarantee that file attributes or metadata necessary to access the file are committed. The manual notes that, if metadata have already been committed, fdatasync() can be more efficient than fsync(). Do not silently import stronger or different wording from another operating system’s man page. For a program that requires a file to be in a known state, FreeBSD’s manual specifically points to fsync() as the conservative choice.
Both calls may fail. The manual lists invalid descriptors, use of a socket descriptor, and storage errors among failure conditions. FreeBSD’s documentation also warns that fdatasync() does not currently guarantee completion of enqueued AIO requests for the file before it returns. If the application uses asynchronous I/O, combine the completion protocol with synchronization rather than assuming that a call on the descriptor reaps outstanding work.
The smaller sync(2) system call is not a substitute for an application-level commit. The manual describes it as forcing dirty buffers toward disk and states that it may return before the buffers are completely flushed. It is periodically issued by the kernel’s syncer process. Use a file-specific operation when a particular file descriptor is the consistency boundary.
Handle errors as durability failures
A process that ignores fsync() errors can report success for data that the operating system could not synchronize. A robust writer propagates failures into its transaction result, emits an actionable log entry, and avoids acknowledging a durable commit to a remote client until the intended boundary has succeeded. An error does not always tell the application how much of the operation reached storage; do not assume that failure means “nothing changed.”
The minimal C shape is:
#include <errno.h>
#include <unistd.h>
int
commit_file(int fd)
{
if (fsync(fd) == -1)
return -1; /* caller must preserve errno and fail the commit */
return 0;
}
This only illustrates error handling around the system call. The caller must still verify that all intended writes completed first, decide what to do if close reports an error, and define whether retries are safe. Repeating a write after an ambiguous failure can duplicate a logical operation unless the data format provides an idempotency key or sequence number.
For a user-level diagnostic on a named regular file, FreeBSD provides an fsync utility that invokes fsync(2) for each path:
fsync /var/log/example.log
status=$?
if [ "$status" -ne 0 ]; then
echo "fsync utility reported an error" >&2
exit "$status"
fi
The command is a manual aid, not a crash-consistency test. It acts on files by path, and its error behavior should be checked against the installed fsync(1) manual. It cannot flush application-private buffers that have not been written, and it cannot make a multi-file application protocol atomic.
Distinguish file synchronization from namespace transactions
A file’s contents and the directory entry that names it are separate concerns. Programs that create a new file, replace a file by rename, or update a manifest often need a carefully ordered protocol for both content and namespace state. Do not assume that synchronizing the file descriptor alone guarantees that the rename or containing directory entry will survive every crash model. The FreeBSD fsync(2) contract above is about the referenced file; verify the documented behavior for the filesystem and exact operation you rely on.
Likewise, atomic visibility of a rename to concurrent processes does not automatically prove crash durability of the entire update. A database may require a write-ahead log, barriers, checksums, or a storage engine protocol beyond simple fsync() calls. Follow the application’s documented durability mode. Do not add ad hoc sync calls to a database’s data directory while its service is active unless its vendor procedure explicitly permits it.
For UFS, soft updates and journaling address filesystem metadata consistency and recovery behavior, while application-level commit semantics remain a separate layer. A filesystem that can recover its structure after a crash does not know whether the last business transaction was complete. Conversely, application-level logging cannot compensate for hardware that acknowledges writes before they are safe under the required failure model.
Device caches and virtualized storage
“Permanent storage” is interpreted across the operating system, filesystem, device driver, controller, transport, hypervisor, and storage device. Volatile write-back caches, battery-backed caches, virtual disk implementations, and remote storage may have different semantics. The application can only rely on documented flush propagation and honest completion at each layer. A controller with protected cache can offer different guarantees from an unprotected drive cache, but the configuration must be verified rather than assumed.
Before a production durability claim, record the filesystem and mount options, storage topology, controller firmware and cache policy, virtual-machine or cloud volume contract, and the tested failure model. A controlled power-cut or fault-injection test is only appropriate in a disposable lab with recovery procedures; never create a destructive outage merely to test a production volume. Use a dedicated test dataset or virtual disk and verify checksums and application-level records after restart.
Record the storage context with each benchmark: filesystem, mount, device path, logical and physical sector sizes if available, queueing configuration, guest/host boundary, and cache policy. A virtual machine may report a completed flush because its virtual device accepted it; the full guarantee then depends on how the host maps that request to the underlying device. Ask the platform provider for the volume’s durability contract rather than extrapolating from a local disk test.
Measure synchronization cost instead of removing it blindly
Frequent synchronization can add latency, especially if each call forces a physical or remote completion. Batching records within a bounded interval or committing a group can reduce overhead, but increases the amount of acknowledged work that can be lost on failure. That is an application-level tradeoff: quantify the maximum uncommitted interval and make it visible in the service’s durability contract.
Collect latency percentiles for synchronization calls under representative load, correlate them with storage queue depth and device errors, and test a workload with the same block sizes and concurrency as production. A benchmark that writes to memory-backed storage or a disposable cache configuration may say little about the deployed path. Do not remove fsync() only because a synthetic benchmark becomes faster; first establish whether that changes the acceptable loss window.
Measure tail latency, not just average throughput. A sync call can block behind unrelated I/O or a device queue, and group commit behavior in the application may change which request pays the cost. Keep a separate counter for synchronization failures and latency histograms at the application layer. If a service is slow, distinguish time spent formatting, writing, waiting for completion, and retrying an error; a single write-latency number can hide the actual bottleneck.
Acceptance criteria
For each durable update, document which process buffer is flushed, which file descriptor is synchronized, whether the data spans a rename or multiple files, how synchronization errors affect acknowledgments, and what storage stack lies below the filesystem. Test recovery with checksums or application-level transaction IDs on a nonproduction fixture. Confirm that the code does not confuse stdio flushing, sync(2), filesystem recovery, and transaction commit.
Choose the narrowest interface that satisfies a written recovery objective, preserve every error, and state the limits of what was actually verified. “We call fsync” is not by itself a complete durability design.
Related:
- UFS dump and restore on FreeBSD: Incrementals, Snapshots, and Recovery
- FreeBSD newfs: Provision UFS Filesystems Deliberately
Sources: