XFS Online Scrub: Metadata Validation, Repair Boundaries, and Recovery
Run XFS online scrub safely by understanding metadata cross-checks, repair limits, data scans, kernel support, backups, and the offline xfs_repair boundary.
XFS online scrub is a mounted-filesystem consistency check that asks the kernel to validate metadata while the filesystem remains in service. The xfs_scrub userspace program coordinates work, interprets results, and may request supported repairs. This is valuable for detecting corruption without taking a volume offline, but it is not an always-on guarantee, a backup, or a universal replacement for offline xfs_repair.
The word “scrub” can also refer to reading file data to detect media errors. XFS metadata validation and a full read of every file extent are different workloads. The latter can consume substantial I/O and is explicitly selected by an option in xfs_scrub; do not assume an ordinary metadata run has read every user-data block. A safe operating plan defines what is being checked, what may be repaired, and how to proceed when repair is unsupported or incomplete.
How online metadata checks work
The kernel scrub interface examines metadata objects under the filesystem’s normal locking and resource rules. It checks values and cross-references related structures. XFS stores redundant information for some metadata, so the kernel can sometimes reconstruct one structure from another known-good copy. Online repair is therefore bounded by available redundancy and by the repairs implemented in the running kernel.
The userspace program schedules scrub work in phases because some checks depend on earlier repairs or verified metadata. When it finds a problem, it reports the result. If repair is enabled and the kernel supports the required operation, the tool can ask the kernel to rebuild supported structures and then re-check them. A successful repair log is useful evidence, but it does not prove every byte of file content is correct or that a separate storage device is healthy.
The feature requires a kernel new enough to expose the XFS metadata-scrub interface and an xfsprogs build with the relevant tool support. Distributions backport features and package versions independently, so use the installed xfs_scrub -V, kernel release, and local man page rather than a generic version-number guess. A mountable XFS filesystem does not necessarily imply that every online scrub/repair operation is supported.
Start with a non-repairing inspection
Identify the actual mountpoint and backing source first. Use a mountpoint, not a guessed block-device path, for xfs_scrub. The -n option requests metadata checking without repair or optimization. It is a sensible first diagnostic run, but even read-only checking can consume CPU, memory, locks, and I/O. Schedule it with awareness of production latency and filesystem size.
findmnt -no TARGET,SOURCE,FSTYPE,OPTIONS --target /srv/data
xfs_scrub -V
sudo xfs_scrub -n -v /srv/data
The final command is intentionally read-only with respect to repairs and optimizations, but run it only after confirming /srv/data is the intended mounted XFS filesystem. Preserve stdout, stderr, exit status, kernel messages, kernel version, xfsprogs version, and mount options. xfs_scrub can return a nonzero status for filesystem findings, optimization opportunities, or operational errors; interpret the documented exit-code bits rather than treating every nonzero result as the same failure.
Do not infer “clean” from a short run without a clear status summary. Confirm that the run completed all phases and that the tool did not stop due to a kernel capability or I/O error. Read the installed manual for version-specific output and flags. If the result says a repair is required, preserve the evidence and assess backups and device health before authorizing a modifying run.
Repair is conditional, not magical
Without -n, xfs_scrub can ask the kernel to make repairs and optimizations supported by the current kernel. Those operations are not equivalent to arbitrary offline reconstruction: the kernel repair framework generally uses available primary and secondary metadata and is constrained by live-filesystem locking. Some corruption cannot be repaired online. If a problem remains, the documented recovery path is to unmount the filesystem and run xfs_repair according to its manual and the site’s recovery procedure.
Never run xfs_repair against a mounted filesystem. It is an offline repair tool and can cause severe damage if used on an active volume. Do not pass -L to clear a log as a routine way to make a mount succeed; discarding the log can lose metadata updates and should be considered only under a reviewed last-resort recovery plan with a verified backup.
An online repair may involve trimming free space or other optimization behavior depending on tool options and filesystem properties. Those operations can generate I/O and alter the storage state even if they are not “repairing corruption.” Inspect the installed xfs_scrub(8) options and site policy before scheduling a normal repair-enabled run. If a maintenance system runs scrub automatically, understand whether it is in check, optimize, or repair mode.
Metadata checks and file-data reads differ
The -x option to xfs_scrub requests reading file data extents to look for disk errors. This can issue direct reads to the block device and, for some SCSI devices, use READ VERIFY commands. It may take a long time and create a substantial I/O load. It is not necessary to include it in every metadata check, and its absence means the run did not verify every data extent by reading it.
Use a data scan only when the desired coverage and storage impact are understood. Check the block device, multipath or RAID layer, device error counters, storage latency, and application SLOs. A media error may identify a disk offset and affected inode, but the remediation still depends on redundancy, backup, and hardware replacement policies. Filesystem metadata repair cannot restore user data that exists nowhere else.
Likewise, a clean scrub is point-in-time evidence. New corruption can occur later, latent media failures can appear after the scan, and online scrub cannot substitute for restore testing. Keep independent backups and verify that the backup can be mounted or restored. Scrub supports detection and some self-repair; it is one control in a storage integrity program.
Read results alongside kernel and device evidence
When scrub reports a problem, preserve the exact diagnostic text and correlate it with dmesg, journal logs, device-mapper state, RAID status, NVMe/SCSI error counters, and recent resets. A transport error during scrub may be a lower-layer failure rather than XFS metadata inconsistency. Conversely, repeated XFS metadata reports on a stable device need filesystem-specific investigation.
Record whether the filesystem was mounted read-write, whether snapshots or reflinks are in use, the current workload, and the duration of each phase. XFS allocation groups distribute metadata work; a localized problem may not correlate with one user directory. Do not delete “bad” files based on a scrub report until the affected inode and data have been backed up and the report’s exact meaning is understood.
If scrub activity causes unacceptable latency, stop or reschedule it according to the tool’s documented controls. Do not kill repair work without understanding whether a kernel operation is in flight. Use the background mode or service policy provided by the installed xfsprogs package only after reviewing its priority and resource limits; package defaults differ.
A cautious repair workflow
- Verify the mountpoint, source UUID, filesystem type, kernel, xfsprogs version, mount options, and recent storage errors.
- Confirm a restorable backup and establish a maintenance window appropriate to the filesystem size and workload.
- Run
xfs_scrub -n -v MOUNTPOINTand save output, logs, and exit status. - Classify each finding as corruption, optimization, unsupported repair, or operational failure. Check device health before repeating a scan.
- If online repair is approved, use the documented xfsprogs invocation and monitor application latency, I/O, and kernel logs.
- Re-run a non-repairing check to confirm the result; if the issue remains or online repair is unsupported, plan an unmount and offline recovery with
xfs_repair.
This workflow deliberately separates diagnosis from modification. A symptom such as “filesystem mounted read-only” can be caused by I/O errors or forced shutdown, and repeated online repairs will not fix a failing device. Ensure the storage path is stable before asking the filesystem to rebuild metadata.
Acceptance criteria and maintenance cadence
Define which filesystems are included, how frequently scrub runs, which mode is used, and what alerts on uncorrected errors or operational failures. Record completion time, phase progress, exit status, error count, and any repair action. The monitoring system should distinguish “no corruption found” from “scan could not complete.” Include a controlled test filesystem to verify that the installed kernel and package support the expected operations.
For service-level acceptance, demonstrate that routine metadata checks complete within the maintenance window without breaching storage latency objectives. If data scans are required, measure them separately. Verify recovery by restoring from backup and by following the unmounted xfs_repair procedure in a disposable copy or lab image. Do not learn a destructive recovery flag during a live incident.
The production claim should be limited and precise: “this supported online scrub run checked the metadata objects exposed by this kernel and reported these findings.” It should not say “all XFS data is verified” unless the data-scan path actually read the extents, and it should not say “the filesystem is guaranteed healthy.” Preserve the reports as maintenance evidence and keep the offline repair and restore paths current.
Related:
- Fixing a Corrupted ext4 or XFS Filesystem with fsck
- Btrfs Copy-on-Write and Snapshots: What Is Shared, What Changes, and What Can Fail
Sources: