Skip to content
LinuxFix Published Updated 6 min readViews unavailable

Fixing a Corrupted ext4 or XFS Filesystem with fsck

The correct order of operations for repairing a corrupted Linux filesystem, and what to do when the repair tool reports it can't fix something automatically.

Filesystem corruption on Linux - usually surfacing as mount failures, unexpected read-only remounts, or explicit kernel errors about filesystem inconsistency in dmesg - has a correct, sequential repair process, and the single most common way to make it worse is skipping straight to running a repair tool without first getting the filesystem into a safe state to be repaired.

Step one: unmount it, or boot from external media

Never run fsck (for ext4) or xfs_repair (for XFS) against a mounted, actively-in-use filesystem. Both tools assume exclusive access to the filesystem’s on-disk structures while repairing them, and a filesystem being simultaneously modified by running processes while a repair tool is rewriting its metadata is a recipe for making corruption meaningfully worse rather than fixing it.

For the root filesystem specifically, this means booting from a live/rescue environment (a USB installer’s rescue mode, or a distribution-specific rescue target) rather than attempting to repair root while it’s mounted and in active use by the running system:

umount /dev/sdb1

for a non-root filesystem you can simply unmount, or boot into rescue mode for anything that can’t be unmounted while running.

Confirm the exact block device before a write repair

A repair tool cannot protect data if it is pointed at the wrong volume. Before a command that may change metadata, map the affected mount point to its source and filesystem type, then verify that source in the block-device tree:

lsblk --fs -o NAME,PATH,FSTYPE,UUID,LABEL,SIZE,MOUNTPOINTS
findmnt --target /mnt/data --output SOURCE,FSTYPE,TARGET

Replace /mnt/data only with the affected mount point, and compare the reported UUID, type, and target with the incident evidence before unmounting. Do not reuse an example /dev/sdX name from another machine: enumeration can differ between boots and hosts. On encrypted, RAID, device-mapper, or LVM layouts, inspect the stacked relationships and identify the mapped device that carries the filesystem instead of blindly repairing a component disk or lower layer. Avoid fsck -A as a generic response to one failing volume; it selects filesystems from /etc/fstab and dispatches to filesystem-specific checkers, so it can inspect more than the incident target. If the device reports I/O errors or the data is valuable, preserve a verified image or backup before any modifying repair attempt.

Reading dmesg first, to understand what kind of corruption you’re dealing with

Before running a repair tool, dmesg (or the last boot’s kernel log) usually contains the actual filesystem-level error that triggered the problem, which meaningfully informs what to expect from repair:

dmesg | grep -iE 'ext4|xfs|EXT4-fs|XFS'

A message about a specific inode, directory entry, or journal replay failure gives you a much more informed sense of the corruption’s scope before you commit to a repair pass, versus a vaguer “unable to mount” error that could indicate anything from trivial journal inconsistency to serious structural damage.

ext4: fsck in stages, not blindly with -y

For ext4, fsck.ext4 (invoked via the generic fsck wrapper) should generally be run first WITHOUT the automatic-yes flag, so you can see what it proposes to fix before approving repairs:

fsck.ext4 -f /dev/sdb1

The -f forces checking even when the filesystem appears clean; an unclean shutdown alone does not prove that corruption exists, since journaled ext4 normally replays committed transactions during recovery. Use an offline check when logs, mount failures, or other evidence justify it, and review proposed fixes on important data before blanket-approving them with -y. Some repairs, such as removing an unrecoverable inode, can lose data and should not be approved reflexively.

Once you understand the scope, -y automates approval for a full repair pass:

fsck.ext4 -fy /dev/sdb1

XFS: a fundamentally different repair philosophy

XFS’s repair tool, xfs_repair, works differently from fsck.ext4 in a way that matters for how you approach it. xfs_repair does not replay a dirty XFS log. The kernel replays that log when the filesystem is mounted; if xfs_repair detects a dirty log, it exits and instructs the operator to mount and cleanly unmount the filesystem on a compatible machine before retrying:

xfs_repair /dev/sdb1

If mounting fails and log replay cannot be completed, xfs_repair -L can force the log to be zeroed. This discards pending metadata updates and can cause significant filesystem damage, so it is a last resort after preserving the device or image and reviewing the recovery options:

xfs_repair -L /dev/sdb1

-L is not a routine first repair attempt. Do not use it unless the normal kernel log-replay path cannot recover the filesystem and the data-loss consequences are understood.

When the repair tool reports unrecoverable corruption

If fsck.ext4 or xfs_repair completes but reports it moved damaged files into lost+found, or reports data loss it couldn’t avoid, the corruption exceeded what automated repair could reconstruct - this is the tool being honest about a limitation, not a sign it did something wrong. At this point, the practical path is: check lost+found for recovered file fragments (often named by inode number rather than original filename, requiring manual identification based on content), and if the affected data matters and no backup exists, this may be the point to stop attempting further filesystem-level repair and consider whether specialized data-recovery tooling working at the block level (rather than the filesystem-structure level fsck/xfs_repair operate at) has anything further to offer - a meaningfully more involved and less certain process than a straightforward fsck repair.

Confirming the fix and remounting

After a successful repair, re-run the check tool once more in read-only/check-only mode (without the repair flag) to confirm no further inconsistencies are reported before remounting for normal use:

fsck.ext4 -n /dev/sdb1

A clean second pass is meaningful confirmation the repair actually completed successfully, rather than assuming a single repair invocation that didn’t error out necessarily means the filesystem is now fully consistent - some corruption scenarios require more than one repair pass to fully resolve, and verifying with a fresh, non-modifying check is worth the extra few minutes before trusting the filesystem back into production use.

Why this whole sequence matters more than any individual command

The unmount-first, understand-the-corruption-second, repair-in-the-tool’s-intended-order-third sequence exists because filesystem repair tools are working with genuinely dangerous, all-or-nothing operations on structures that, if handled incorrectly, can turn recoverable corruption into unrecoverable data loss. Treating fsck/xfs_repair as “the one command that fixes filesystem problems” without respecting the order of operations they’re actually designed around is the most common way a filesystem repair situation gets meaningfully worse instead of better.

Previewing what a repair would do before committing to it

fsck.ext4 -n /dev/sdb1
xfs_repair -n /dev/sdb1

Both tools have a no-modify/check-only mode, but that is not a substitute for an image of a failing device or a guarantee that every inconsistency will be detected. In particular, xfs_repair -n can miss some metadata-map problems and can report repeated warnings because it cannot fix issues as it encounters them. Keep the target unmounted, preserve important data first, and read the installed tool’s manual for the limits of its check mode.

Why XFS’s log-replay-first design exists

XFS’s log-replay-first recovery is a kernel mount operation, not an action performed by xfs_repair. The distinction is operationally important: normal recovery is to mount and cleanly unmount the filesystem so the kernel can replay the log, then run xfs_repair only if further offline repair is warranted. If the log cannot be replayed by mounting, zeroing it discards pending updates and may lose files or metadata. Do not describe xfs_repair itself as replaying the log.

Related:

Sources:

Comments