FreeBSD SMART Storage Health Operations: Inspect Devices and Schedule Tests
Use smartmontools on FreeBSD to inspect drive health, interpret logs cautiously, run self-tests, monitor devices, and escalate evidence safely.
SMART telemetry can provide useful evidence about a drive’s reported health, error history, and self-test results. It is not a universal prediction system: attributes vary by device, a passing status does not guarantee future reliability, and a controller or bridge may hide the drive’s management interface. Treat smartctl output as one input to storage operations, alongside FreeBSD’s CAM and GEOM views, device logs, application symptoms, and tested backups.
FreeBSD’s base system provides storage discovery tools, while smartctl and smartd are supplied by the smartmontools package. The project supports ATA/SATA, SCSI/SAS, and NVMe devices on multiple operating systems, but exact commands and exposed fields depend on the drive, controller, transport, and smartmontools version. Never translate a manufacturer’s attribute number or raw value into a failure percentage without device-specific authoritative documentation.
Identify devices without guessing
Capture the installed release, package version, storage topology, and stable drive identity before running queries:
freebsd-version -kru
camcontrol devlist -v
geom disk list
pkg info smartmontools
smartctl --scan-open
Install the package if needed using the system’s approved package repository:
pkg install smartmontools
Do not assume da0 refers to the same physical disk after a reboot or hardware change. Correlate device nodes with serial numbers, WWN, enclosure slots, controller paths, and the storage array’s inventory. If a disk sits behind a RAID controller, USB bridge, or virtual machine layer, determine whether that layer exposes SMART commands and which device-type option the smartctl documentation requires.
Start with a single, explicitly identified device:
smartctl -i /dev/da0
smartctl -H /dev/da0
smartctl -x /dev/da0
Use only the actual device path and preserve the output. The -i option reports identification information, -H requests the device health status, and -x requests broader SMART and non-SMART data in current smartmontools documentation. ATA, SCSI/SAS, and NVMe output differs. Consult smartctl(8) for the installed package version and avoid assuming a field shown for one protocol exists for another.
Read health results as evidence, not a verdict
A device-reported overall health result is a summary, not a certificate that the drive is safe. A PASSED result does not rule out a sudden mechanical, electrical, firmware, transport, or controller failure. A failing result deserves prompt escalation, but immediate replacement still needs to follow the array, filesystem, workload, and redundancy procedure.
Review error logs, self-test logs, temperature, media errors, interface counters, wear indicators, and device identity where supported. Attributes may be vendor-specific, normalized, thresholded, or encoded in raw fields. Compare a device’s values over time and with the manufacturer’s documentation. Do not apply generic thresholds such as a universal reallocated-sector number across all drive families.
Check host logs for CAM errors, resets, timeouts, or transport faults. A drive can report normal SMART state while its cable, HBA, power, enclosure, or SAS path is unstable. Conversely, an error counter can persist from a past event even after a component has been replaced. Preserve the timestamp and baseline before clearing or resetting anything.
When a device is behind hardware RAID, direct SMART access may be blocked or require a controller-specific device selector. Read the smartmontools device-type documentation and controller manual. Do not guess a numeric disk index or issue a write-capable query through a controller without understanding its mapping. Keep the array’s own health report as a separate source of truth.
Run self-tests in a maintenance window
SMART self-tests are initiated on the device and can consume time and I/O. They are not a substitute for a ZFS scrub, UFS check, backup verification, or a vendor qualification test. Read the drive and smartctl documentation before running a test through a shared controller or during an already degraded array.
Begin with a short test on a noncritical or approved device:
smartctl -t short /dev/da0
smartctl -l selftest /dev/da0
The first command starts a device self-test where supported; the second reads the test log. Consult the output for an estimated completion time, then query again after that interval. A self-test may continue after the smartctl process exits. Do not repeatedly launch tests because the log appears unchanged; verify whether one is already running.
An extended test can take much longer and may affect latency. Schedule it one drive at a time when the storage design requires a maintenance window. Do not test every member of a degraded mirror or RAID-Z at once. Record the exact device serial and the array member mapping so a test result is not attributed to the wrong drive.
Some drives support captive versus background tests, selective ranges, abort commands, or protocol-specific behavior. Do not use these modes without checking smartctl(8) and the device manual. A command accepted by a bridge may not be delivered to the target device, and a test’s completion status must be interpreted using its log result.
Interpret exit status and automation carefully
smartctl uses an exit-status bitmask to report different conditions, not a simple binary “healthy/unhealthy” code. Scripts should preserve the raw output and status and decode the bits according to smartctl(8) for the installed version. Do not treat every nonzero code as a transport failure or suppress all warnings by adding a permissive option.
For periodic monitoring, smartd can monitor supported devices and report selected conditions. Establish which device paths are stable, which attributes and logs matter for those devices, how notifications are delivered, and who owns an alert. Validate mail or notification transport before relying on it. A daemon configured to send alerts through a broken system-mail path creates false confidence.
Monitoring cadence should account for device workload and service policy. Frequent polling is not necessarily better and may interact poorly with power-saving devices, bridges, or controller behavior. The smartctl documentation includes a power-mode option to avoid waking some ATA/SCSI devices; use such options only after confirming the behavior for the exact device type. Do not accidentally spin up a cold-storage disk with a monitoring query if that violates the access plan.
Triage a warning or a missing device
If smartctl cannot open a device, first verify the path, permissions, package installation, and whether the kernel sees the device:
camcontrol devlist -v
geom disk list
dmesg | tail -100
smartctl --scan-open
If the host sees the disk but smartctl cannot read it, investigate the driver, pass-through support, controller mode, bridge chipset, and correct device type. Avoid immediately using -d sat, -d scsi, or a controller-specific selector; determine the correct protocol from the hardware documentation and scan output.
If the health summary fails or a self-test reports an error, preserve the full output, serial, error log, test timestamp, pool status, and host messages. Notify the storage owner and follow the approved replace or failover process. Do not pull or offline a drive based only on a generic guide if it participates in a redundant vdev or hardware RAID group.
If host I/O errors appear but SMART data is clear, troubleshoot the path: inspect HBA and enclosure logs, cables, link resets, power events, CAM state, and array status. If SMART attributes change but no host errors are visible, continue trend collection while planning a vendor-supported action. Neither condition should be dismissed because the filesystem remains mounted.
Separate device monitoring from data protection
SMART does not protect data. It cannot reconstruct a failed disk, restore a deleted file, validate application consistency, or prove a backup can be restored. ZFS checksums and redundancy provide different integrity and repair behavior; ZFS scrub and resilver operations should follow their own procedures. For UFS or other filesystems, follow their appropriate checks and recovery tools.
Maintain verified, independent backups and periodically perform test restores. When a SMART warning leads to replacement, map the physical drive to the logical pool or array member before service. Track resilver or rebuild progress, keep another copy available, and perform a follow-up integrity check after hardware remediation.
Do not use a SMART self-test as a destructive burn-in or secure erase. Long tests are diagnostic operations; they are not a comprehensive guarantee of remaining life. For asset disposition, follow a separate approved sanitization process.
Acceptance criteria and evidence retention
For a routine inspection, acceptance means the intended physical device was queried through a known path, the output and exit status were captured, the current self-test state is understood, and the result is mapped to the correct storage member. For a warning investigation, acceptance requires an explicit disposition: continue monitoring, schedule replacement, escalate to the vendor, or close with documented evidence.
Record FreeBSD and smartmontools versions, device protocol, model, serial or approved identifier, controller and transport, query time, command line, output, self-test result, and next action. Redact serials when sharing externally if asset policy requires it. Keep only the data needed to diagnose and trend the device.
Related:
- FreeBSD CAM Storage Diagnostics: Tracing Devices from Bus to GEOM
- FreeBSD GEOM I/O Observability: Read gstat and iostat Without Double Counting
Sources: