FreeBSD ZFS Scrub and Resilver Operations: Verify Data and Restore Redundancy
Operate ZFS scrubs and resilvers on FreeBSD with safe scheduling, status interpretation, error triage, workload controls, and measurable recovery criteria.
A ZFS scrub and a resilver both read pool data, but they answer different operational questions. A scrub checks the data that ZFS can reach against stored checksums and, when redundant copies exist, can repair a damaged copy from a valid one. A resilver reconstructs the portions ZFS knows are stale after a device is attached, replaced, or otherwise made current. Neither operation is a backup, and neither makes a pool with missing redundancy safe from a second failure.
This runbook treats these operations as controlled maintenance: establish which pool and devices are involved, capture the starting state, run only the appropriate operation, observe its effect, then record a completion result. Commands below are examples for a pool named tank. Replace names only after confirming them locally.
Establish a trustworthy baseline
Before starting work, identify the pool, topology, device paths, and any existing activity. Do not infer the pool layout from device names alone. A pool may include mirrors, RAID-Z, special vdevs, cache devices, logs, or multiple top-level vdevs, and loss characteristics differ across them.
zpool list -v
zpool status -v tank
zpool get health tank
zpool iostat -v tank 1 5
Keep a timestamped copy of the status output with the maintenance ticket. The status tree identifies the pool state and child vdev states; READ, WRITE, and CKSUM columns are cumulative counters, not a latency dashboard. The errors section is important: “No known data errors” means no currently recorded uncorrectable data errors, not that the pool is backed up or that every application-level file is valid.
Confirm recent backups by checking the destination, retention, and a restore path, not by seeing a successful cron mail alone. If there is an active resilver, do not initiate a scrub to “help”; ZFS permits only one scrub or resilver operation per pool at a time. If the pool is already reporting degraded or faulted devices, first determine whether an operator is replacing hardware or recovering a missing device. A scrub is not a substitute for resolving a device fault.
Run a scrub when the question is integrity
A regular scrub traverses allocated pool data and verifies checksums. On mirror, RAID-Z, or dRAID vdevs, ZFS can repair damage when it has a good redundant copy. A single-copy pool can detect checksum mismatch but cannot reconstruct a correct block from a copy that does not exist. Scrubbing is I/O-intensive, and only one such operation can run per pool at once.
Start a scrub during an agreed maintenance window:
zpool scrub tank
zpool status tank
Track it with periodic status checks rather than repeatedly restarting it. The scan line reports scanned and issued bytes, rate, repaired bytes, completion percentage, elapsed time, and estimate when available. Estimates are not service-level guarantees: workload, errors, device retries, pool writes, and system uptime affect the remaining estimate. Live writes can also make progress exceed 100 percent briefly.
To pause a scrub because foreground service has priority, use:
zpool scrub -p tank
The pause state and checkpoint are persisted periodically. A later zpool scrub tank resumes a paused scrub from its saved checkpoint. Use zpool scrub -s tank only when the operation should be stopped rather than paused. Stopping is a deliberate cancellation; do not confuse it with a temporary maintenance pause. For a script that should wait, zpool scrub -w tank blocks until completion. Prefer an external job timeout and alerting policy over an unbounded shell session.
FreeBSD periodic can schedule scrubs by pool and last-scrub age. Review the current periodic.conf(5) settings before enabling it because a system can have pools with different workloads and maintenance windows. A monthly scrub is a common operational target documented by the Handbook, but fleet cadence should account for data criticality, device scale, workload, and available maintenance time. Do not schedule every pool simultaneously without considering I/O pressure and memory use.
Understand resilvering before replacing a device
A resilver is normally triggered by a ZFS topology change such as attaching or replacing a device. It uses ZFS’s knowledge of which data is out of date and writes the needed copies to the new or returning device. It is not a complete checksum audit of every block, so a routine scrub still has a separate purpose.
For an approved replacement, inspect the current topology and map the failed member to the physical drive by serial number or enclosure slot. Never remove a disk based only on adaN or daN ordering if device enumeration may change after reboot. The exact zpool replace syntax depends on the pool’s existing leaf path and the replacement device label. Read zpool-replace(8) for the local release and copy the current leaf exactly; this article intentionally does not provide a universal destructive replacement command.
After the approved change:
zpool status -v tank
zpool iostat -v tank 1 10
Confirm that the expected leaf is ONLINE or is actively resilvering, that the destination is the intended device, and that the pool has not acquired new errors. Wait for the scan to finish and capture the final status. Do not manually start zpool resilver tank as a routine “continue” action: the command restarts an already running resilver from the beginning. Automatic resilvering is normal after a correct replace operation; a manual resilver is for a deliberate and documented case.
During a resilver, avoid unrelated topology changes, firmware experiments, and untested power cycling. Keep workload within a known operating envelope, but do not silently throttle an operation that is protecting a degraded pool without reviewing risk. If the system must reboot, preserve evidence and consult the status after import. The scan is restartable, but repeated interruptions extend the vulnerable period.
Triage scan results without erasing evidence
At completion, record the scan summary, pool topology, device counters, and errors section. A nonzero repaired-byte value during a scrub means ZFS repaired data from redundancy; it is not a reason to clear all counters and close the incident. Investigate whether errors recur, correlate with a disk, cable, HBA, enclosure, power, or firmware event, and verify application data from backup when business impact warrants it.
If zpool status -v lists permanent errors, it identifies data that could not be repaired and may show affected paths. Checksum counters on a leaf can indicate errors seen by ZFS, but counters alone do not distinguish media defects from transport faults. Correlate with system logs, SMART or vendor telemetry where appropriate, CAM messages, and a controlled replacement plan. Do not run zpool clear before saving status and diagnostics. Clearing resets recorded error counters; it does not repair data or prove that a physical fault is gone.
If a scrub repaired damage, schedule a follow-up observation or scrub according to the incident policy. A clean second scan after hardware correction is useful evidence, but it is not a substitute for restoring affected files or validating application checksums. For encrypted datasets, keys need not be loaded for a scrub to verify pool blocks, but the verbose status may be unable to name an unrepairable file if the key is unavailable. This distinction matters when preparing the recovery environment.
Observe impact and define completion
Use zpool iostat -v with an interval to see per-vdev throughput and queue behavior while scans run. Compare the maintenance period with a baseline from similar workload hours. A throughput number by itself does not establish that service impact is acceptable; check application latency, queue depth, and workload-specific error budgets. If foreground I/O is disrupted, use the documented pause behavior or reschedule instead of repeatedly canceling and restarting.
For a scrub, completion means status reports a completed scrub with the intended pool, final repaired-byte count, and no unexplained permanent errors. For a resilver, completion means the new or returning leaf is healthy, the scan is finished, topology redundancy is restored, and the pool remains stable afterward. Record who approved the change, exact devices, start and end times, any observed errors, and follow-up work.
Keep separate evidence for backup restore tests, application-level consistency checks, and pool integrity. ZFS checksums and redundancy address block integrity and reconstruction within the pool; they do not protect against deletion, ransomware, administrator error, loss of the entire pool, or application-level corruption replicated consistently to every copy.
Related:
- Recovering a ZFS Pool That Won’t Import on FreeBSD
- How to Replicate FreeBSD ZFS Datasets Safely with send and receive
Sources: