ext4 Journal Semantics: Ordered Data, Recovery, and Application Durability
Separate ext4 metadata recovery from application durability by tracing JBD2 transactions, data modes, fsync boundaries, rename, and crash testing.
“The filesystem is journaled” is not a complete durability guarantee. A journal can help ext4 recover a consistent metadata structure after an interrupted update, but that does not mean an application record is atomic, that the latest write reached stable media, or that a rename survives a sudden power loss. Those properties depend on the filesystem’s data mode, the application’s synchronization protocol, the block stack, and the device’s handling of flushes.
This distinction matters for databases, package managers, editors that replace files, and services that checkpoint state. The right mental model has at least three layers: ext4 and JBD2 transaction recovery; VFS and system-call completion; and the storage device’s volatile cache and persistence behavior. A successful write() normally says that bytes were accepted into the kernel’s buffered I/O path, not that a disk platter or nonvolatile flash has committed them.
What JBD2 journals and what it does not
JBD2 records filesystem updates in transactions and allows ext4 to replay committed journal state after an unclean shutdown. In the common metadata-journaling modes, the journal protects filesystem metadata; ordinary file contents are not copied into the journal. Recovery can therefore restore a structurally consistent filesystem without reconstructing the exact application-level contents that existed before a crash.
The data=ordered mode is ext4’s documented default mode when no other mode is selected. It journals metadata and orders relevant data writes before the metadata transaction that exposes them is committed. That ordering reduces a class of stale-data exposure, but it is not a promise that every recently written byte is persistent at every instant. data=writeback journals metadata without the same data-before-metadata ordering guarantee. data=journal writes file data and metadata through the journal, which changes performance and write amplification; it still does not define an application’s multi-file transaction protocol.
Do not infer mount behavior from a generic Linux description alone. Inspect the live filesystem’s type and options with findmnt, mount, and the distribution’s configuration. Kernel, ext4 feature flags, mount options, and the storage stack all matter. The journal’s existence also does not replace backups or offline repair tools when metadata damage exceeds what journal replay can resolve.
Transaction commit is not the same as fsync()
The journal groups metadata operations into transactions, but applications request persistence through interfaces such as fsync() or fdatasync(). These calls ask the kernel to synchronize file data and the metadata needed for the file’s retrieval; they do not automatically synchronize every directory entry or unrelated file. The exact syscall sequence required depends on what the application promises to preserve.
For an atomic replacement pattern, an application commonly creates a temporary file in the target directory, writes the full new contents, calls fsync() on the temporary file, renames it over the old name, then calls fsync() on the parent directory. The file sync protects the new file’s contents and metadata as required by the API. The directory sync is needed when durability of the name-to-inode update matters. Keeping the temporary file on the same filesystem makes the rename atomic with respect to namespace observers; it does not by itself guarantee crash durability.
fdatasync() may omit metadata that is not required for subsequent data retrieval, while fsync() requests file synchronization including metadata. Applications must check errors from writes, syncs, close, and rename. A delayed I/O error can surface at a later call. If the filesystem is mounted over device mapper, RAID, a virtual disk, or network storage, confirm that each layer propagates flush and FUA semantics correctly. A consumer SSD or controller that lies about stable writes can invalidate the expected durability contract.
A small atomic-replacement experiment
The following Python example implements the replacement sequence for a single text file. It intentionally operates only on the path supplied by the caller. Use a disposable directory when testing, inspect the final content, and do not use a production data file as a lab target.
#!/usr/bin/env python3
import os
import pathlib
import sys
import tempfile
target = pathlib.Path(sys.argv[1]).resolve()
parent = target.parent
payload = b"new complete record\n"
fd, temporary_name = tempfile.mkstemp(prefix=f".{target.name}.", dir=parent)
temporary = pathlib.Path(temporary_name)
try:
with os.fdopen(fd, "wb") as stream:
stream.write(payload)
stream.flush()
os.fsync(stream.fileno())
os.replace(temporary, target)
directory_fd = os.open(parent, os.O_RDONLY)
try:
os.fsync(directory_fd)
finally:
os.close(directory_fd)
finally:
temporary.unlink(missing_ok=True)
The temporary file is created in the same directory so the rename stays on one filesystem. If a write or synchronization fails, the program raises an exception; production code should report the error and preserve enough state for recovery. On platforms or filesystems where syncing a directory descriptor is unsupported, the example must be adapted to the documented platform contract rather than silently treating the rename as durable.
This program demonstrates ordering of system calls, not the physical behavior of a particular disk. To make the guarantee meaningful, verify the filesystem, mount options, device cache policy, and any virtualization or RAID layers. An application that updates both a data file and a separate index needs a transactional design above this single-file pattern.
Delayed allocation and rename patterns
ext4’s delayed allocation postpones choosing physical blocks until writeback. This can improve allocation quality, but it means the application must not equate a successful buffered write with completed allocation. The auto_da_alloc behavior exists to detect common replace-via-rename and replace-via-truncate patterns and order delayed allocation to reduce zero-length or stale-content outcomes after a crash. That compatibility behavior is not an application transaction and does not make arbitrary update sequences durable.
When building an update protocol, distinguish the old name, temporary name, inode contents, directory entry, and any metadata record that points at them. For example, an application may correctly sync a new file but forget to sync the directory after rename; after a crash, data may exist while the expected name does not. Conversely, syncing the directory without first syncing the file can preserve the new name while the file contents are not the intended complete payload.
Use a tested sequence from the relevant API documentation and application’s own recovery design. Database engines often use carefully ordered write-ahead logs, checksums, commit records, and barriers; copying just one fsync() call from their code does not reproduce that protocol. Treat the filesystem as one component in an end-to-end persistence chain.
Journal settings and diagnostic evidence
Capture the current state before changing mount options. findmnt -no SOURCE,FSTYPE,OPTIONS /path identifies the mounted source and visible options. tune2fs -l /dev/DEVICE can report on-disk ext4 feature metadata, but querying the wrong underlying device or running tools against a mounted/active volume without understanding their mode is unsafe. Use read-only options for inspection and consult the installed e2fsprogs manual for the exact version.
findmnt -no TARGET,SOURCE,FSTYPE,OPTIONS --target /var/lib/example
for options in /proc/fs/ext4/*/options; do
[ -r "$options" ] && { printf '%s\n' "$options"; cat "$options"; }
done
The second command lists the per-mount ext4 option files exposed under procfs; their directory names and output depend on the running kernel and mounted devices. They show active mount options, not the on-disk feature list printed by tune2fs -l. An empty glob is not evidence that a filesystem has no journal. Prefer the mounted source and kernel documentation, then record dmesg messages associated with ext4 recovery, block errors, and device resets. A normal mount after a crash proves that recovery completed, not that the application committed its most recent logical transaction.
Do not switch to data=journal as a reflexive “safety” fix. It can change performance and application behavior, and the correct setting is workload-specific. First establish whether the defect is a missing application sync, a device flush problem, an ext4 error, or a backup/restore issue. Validate any mount change on a representative disposable filesystem and benchmark both steady-state and crash-recovery paths.
Crash-consistency testing
Power-loss testing is meaningful only in a controlled environment. Use a disposable virtual machine or test block device with snapshots and a reproducible workload. Kill the VM at recorded points around write, file sync, rename, and directory sync. After recovery, check the application’s invariants, not merely that the filesystem mounts. Include repeated power cuts, device-cache configurations representative of production, and the exact kernel and ext4 feature set.
For an atomic file replacement, acceptable states after a crash are usually the complete old version or the complete new version, never a partial mix. The expected state for a database or package transaction may be different, but it should be written down before testing. Include low-space behavior and I/O errors: a successful path does not demonstrate that cleanup and recovery work when fsync() returns EIO or an allocation fails.
Operational acceptance criteria
Document the filesystem source, UUID, mount options, kernel release, ext4 feature flags, device-mapper or RAID stack, and drive-cache policy. Run a write workload that issues the intended sync sequence, reboot cleanly, then repeat crash tests in a disposable environment. Verify application invariants, error reporting, and restore from backup. Monitor ext4 and block-layer errors separately from application failures.
The acceptance statement should be explicit: after a confirmed successful commit operation, which file contents and names must survive a crash? Which error codes cause the application to stop acknowledging writes? How are inconsistent records detected and repaired? Journal replay answers none of these questions on behalf of the application. A production-grade durability claim is a contract spanning the syscall, filesystem, block stack, controller, device, and recovery logic.
Related:
- Understanding the Linux Page Cache and How Writeback Actually Works
- Fixing a Corrupted ext4 or XFS Filesystem with fsck
Sources: