FreeBSD zfsd Operations: Understand Automated ZFS Fault Responses
Inspect how FreeBSD zfsd handles device events, hot spares, autoreplace, error thresholds, and the limits of automatic ZFS recovery.
zfsd(8) is FreeBSD’s userland fault-management daemon for ZFS. It listens to kernel device-control events and attempts a limited set of actions, including activating a configured hot spare or bringing a returning vdev online. It is not a replacement for monitoring, a backup, or a human replacement procedure. A pool can remain degraded, suspended, or unavailable even while the daemon is running correctly.
The most useful operational model is to treat zfsd as an event-driven helper with explicit prerequisites. Pool topology, hot-spare inventory, device identity, physical paths, and the autoreplace property influence what it can do. A service status of running does not prove that a spare exists, that a returned disk is the expected one, or that the pool has finished resilvering.
Inventory the pool before interpreting events
Capture the current pool topology, health, properties, and available spare devices before enabling or changing automatic behavior:
zpool status -v tank
zpool list -v tank
zpool get autoreplace,autoexpand tank
zpool history -il tank | tail -80
Replace tank with the actual pool name. The commands answer different questions: status reports current vdev state and known errors; list -v exposes the hierarchy and capacity view; get reads pool properties; and history helps establish what an administrator previously changed. None proves the physical identity of a disk. Correlate /dev paths, labels, serial numbers, and enclosure mapping before a physical action.
Inventory any hot spares and confirm that the proposed spare is actually part of the pool’s spare configuration. A disk that is merely installed in the chassis is not necessarily a ZFS spare. Confirm its capacity and connectivity, then retain the expected pool topology in a change record. An automatic spare is useful only if it is reliable and if the system can distinguish the intended replacement from unrelated devices.
Know which events zfsd can handle
The FreeBSD manual describes zfsd as listening to devctl(4) events such as I/O errors and disk removals. On a leaf-vdev disappearance, it can activate an available hot spare. If a device later arrives with a matching ZFS label, it can attempt to online the old vdev; after resilvering, an active spare may detach. The daemon can also use a device’s physical path for replacement decisions when the pool’s autoreplace property is enabled.
These behaviors have important boundaries. A device that returns with a recognizable ZFS label is different from a blank replacement disk at the same physical path. Path matching is not a general guarantee that every new disk in that bay is safe to adopt. The manual documents how zfsd reacts to these conditions, but the administrator remains responsible for checking the resulting topology, disk identity, and resilver state.
The manual also documents thresholds for a vdev to be marked faulted or degraded after repeated I/O errors, delayed I/O events, or checksum errors. Those thresholds and tunables are implementation details that can vary by release. Read the zfsd(8) and vdevprops(7) manuals installed on the running system before changing them. Do not copy an old threshold into configuration merely because it appears in a blog post or an older release manual.
Automatic response is not equivalent to continuous availability. If a mirror loses enough members, or a RAID-Z group loses more devices than its redundancy can tolerate, a spare cannot recreate missing data. zfsd can react to device state; it cannot compensate for insufficient redundancy, common-cause failures, a bad controller, or an incorrect replacement.
Configure spares and autoreplace deliberately
Before adding a spare, inspect the exact device and pool topology. Use stable GEOM labels where your storage design supports them, and verify serial number and capacity separately:
gpart show -l
geom disk list
zpool status -v tank
The command that adds a hot spare changes pool configuration and should be performed only after the device has been positively identified. Use the appropriate zpool add ... spare ... syntax from the installed manual and your approved change plan; this article intentionally does not provide a copy-and-paste device path that could select a live data disk.
autoreplace changes how a newly appearing device at a previously used physical path may be treated. Check its current value first. Enabling it globally without understanding the controller’s path stability can cause an unexpected replacement attempt when devices are moved or enumerated differently. Test the intended behavior with the actual HBA, enclosure, and release in a non-production pool before relying on it for unattended recovery.
Service startup policy is also release-specific. Inspect the installed rc script and local manual, then use service zfsd status and the configured zfsd_enable setting to confirm the daemon’s lifecycle. Do not assume that every FreeBSD installation enables it by default, and do not restart storage services during an active fault simply to make a status screen look cleaner.
Observe an incident from event to recovery
When a disk disappears or accumulates errors, gather a timestamped set of observations before making a change:
date -u
service zfsd status
zpool status -v tank
zpool get autoreplace tank
tail -100 /var/log/messages
camcontrol devlist -v
zfsd logs interesting actions to syslog with the daemon facility and a zfsd identity according to its manual. The actual log destination depends on syslog.conf and the host’s logging setup. Search the configured logs for the daemon identity rather than assuming every system writes to the same file.
If zfsd activates a spare, confirm that the expected spare became active and that the pool’s redundancy state is understood. Track resilver progress with zpool status and avoid starting competing maintenance operations until the pool has returned to the planned state. If the old disk returns, do not unplug or reinsert it repeatedly; observe whether the expected vdev is onlined and how the spare transitions after resilvering.
If the daemon does not act, investigate prerequisites instead of repeatedly restarting it. Is the event visible in system logs? Is a hot spare configured and available? Is autoreplace required for the physical-path case? Does the disk expose the same path? Is the pool already suspended or in a state the daemon cannot repair? These questions narrow the failure boundary without turning an automatic-management feature into an uncontrolled experiment.
Avoid unsafe assumptions about hot replacement
Disk replacement procedures depend on the pool topology and whether the failed member remains present. A replacement command suitable for a mirror may not be appropriate for a RAID-Z member. Use the exact zpool replace procedure for the observed state and check the installed zpool(8) documentation. Keep the vdev identity in the command aligned with the pool status output, not a remembered disk number.
Do not remove a hot spare simply because the dashboard shows a healthy pool. A spare may currently be active and carrying resilvered data, or a permanent replacement may not have completed. Conversely, retaining an active spare indefinitely can leave the system with less spare capacity than expected. Review the zpool status tree and history before adding or removing any device.
Threshold tuning also has tradeoffs. A lower error threshold may mark a slow but recoverable device faulty sooner; a higher threshold can keep a failing member in service longer. Delay and checksum events represent different symptoms and should not be collapsed into one generic “disk bad” signal. Tune only from observed workload behavior, vendor specifications, and a tested recovery plan.
Define acceptance criteria for automation
For each production pool, document whether zfsd is enabled, which hot spares are configured, whether autoreplace is enabled, and how events reach operators. Alert on pool degradation independently of daemon health. A complete test in a disposable environment should verify that a simulated device removal produces the expected event, that a spare is activated only when intended, that recovery progress is visible, and that the final topology matches the change record.
The success criterion is not that a fault was automatically hidden. It is that an event was detected, the documented action occurred, data remained accessible within the design’s limits, and an operator can prove the pool returned to the intended redundancy state. Backups and restore tests remain necessary because zfsd cannot recover data that the pool has irretrievably lost.
Related:
- FreeBSD ZFS Scrub and Resilver Operations: Verify Data and Restore Redundancy
- FreeBSD zpool history: Reconstruct ZFS Administrative Changes
Sources: