FreeBSD CAM Storage Diagnostics: Tracing Devices from Bus to GEOM
Diagnose FreeBSD CAM storage paths from transport and LUN discovery through peripheral drivers, GEOM providers, error sense data, and safe recovery.
When a disk disappears, times out, or returns I/O errors, a /dev name alone does not identify the failing layer. FreeBSD storage can involve a controller, transport, CAM’s XPT and SIM layers, a peripheral driver, GEOM transformations, a filesystem, and an application. A useful diagnostic records the path across those layers before changing device state. This avoids treating a cable error as a filesystem problem or a stale GEOM consumer as a failed drive.
The Common Access Method (CAM) framework organizes storage requests around CAM Control Blocks (CCBs). A peripheral driver creates requests; the XPT transport layer dispatches them to a System Interface Module (SIM); the SIM translates them to commands for a controller or device and returns completion status. GEOM commonly issues block I/O through a CAM peripheral, while character-device access can enter the stack through other paths. The exact layers depend on the device and driver.
Map the stack before touching the device
Start with the provider that failed and work both directions. A typical direct-access disk may appear as da0; ATA/SATA devices can appear through the ada driver; a device on the NVMe CAM path may appear as nda0. NVMe can also use the nvd driver, which follows a different path from CAM’s nda. Do not assume that every /dev disk node is managed by CAM or that camcontrol is the universal command for storage.
The cam(4) architecture distinguishes layers: peripheral drivers know how to translate standard requests, XPT schedules and routes CCBs, and SIM drivers connect to the underlying adapter or transport. A SCSI address is usually represented as bus, target, and logical unit number (LUN). A device name such as da2 is a peripheral instance, not the physical identity of a drive. Record the controller, bus/target/LUN, model, serial or namespace identity where available, GEOM provider, and filesystem or pool using it.
Collect a read-only baseline before rescan, reset, detach, or replacement:
freebsd-version -kru
pciconf -lv
camcontrol devlist -v
geom disk list
gpart show
mount
swapinfo -h
These commands answer different questions. pciconf helps identify PCI controllers. camcontrol devlist -v shows devices attached to CAM with transport details. GEOM output shows providers and transformations. Partition and mount listings show whether a provider is in active use. If a ZFS pool consumes the disk, inspect zpool status and the pool’s device paths before considering hardware removal. Preserve the command output and timestamp with kernel messages from the same incident window.
Query the CAM device without writing to it
The camcontrol(8) utility can list devices and peripheral drivers, inspect inquiry data, test readiness, and read capacity. For a confirmed SCSI direct-access peripheral such as da0, start with bounded read-oriented queries:
camcontrol devlist -v
camcontrol periphlist da0
camcontrol inquiry da0
camcontrol readcap da0
camcontrol tur da0
Substitute the actual SCSI device identifier. inquiry requests identity information, readcap asks the device for capacity information, and tur issues a SCSI Test Unit Ready command. These SCSI commands may not be meaningful for ATA, NVMe, or other non-SCSI devices even when their driver uses CAM; use the relevant driver-specific utility and manual instead. Compare output with kernel messages and the exact device manual. For multipath storage, a path can fail while the logical device remains available through another path; evaluate the path manager and provider state before treating one command’s error as total device loss.
Do not begin diagnosis with arbitrary pass-through commands, camcontrol format, mode-page edits, firmware downloads, or controller resets. camcontrol(8) explicitly warns that improper use can cause data loss or system crashes. The pass(4) device exposes direct access to CAM commands and requires careful access control. Destructive or state-changing operations belong in an approved maintenance plan after a verified backup and a documented recovery procedure.
Read errors as layered evidence
Save the full CAM status, SCSI status, and sense data rather than copying only the last human-readable line. Sense key, additional sense code, and additional sense-code qualifier can distinguish conditions such as a device not being ready from a medium error or a unit attention event. They do not independently prove a physical failure: a controller reset, enclosure event, firmware update, path change, or device initialization can alter the status. Correlate repeated errors with timestamps, device identity, bus resets, link state, controller logs, and application I/O.
Useful evidence includes:
| Layer | Evidence |
|---|---|
| PCI/controller | Vendor/device IDs, driver attach messages, firmware, reset counters |
| CAM transport | Bus/target/LUN, XPT path, SIM, rescan and timeout events |
| Device | Inquiry identity, readiness, capacity, sense data, serial or namespace |
| GEOM | Provider name, labels, consumers, mirror or multipath state |
| Filesystem/application | Mount status, pool status, I/O errors, latency, affected requests |
A timeout can arise because a command did not complete before its deadline. Repeated timeouts may indicate device firmware, cabling, controller saturation, transport errors, power management, or a degraded path. Increasing a timeout can hide the evidence and extend application stalls. First determine if the device is receiving commands, whether the controller is resetting, and whether other devices on the same bus show correlated errors.
If only one peripheral is affected while peers on the same SIM continue to operate, investigate device-level readiness, media, firmware, and path-specific cabling. If many devices fail together, investigate the controller, shared transport, power, backplane, and driver. If CAM reports healthy completions but GEOM or the filesystem reports failures, trace the provider chain and filesystem logs instead of repeatedly querying the physical disk.
Device discovery and controlled rescans
For a genuine hot-plug or newly attached device, camcontrol rescan can request a scan of all buses, one bus, or a specific device/LUN. A bus:target identifier defaults to LUN 0; it does not scan every LUN on that target. The camcontrol(8) manual explicitly says that scanning all LUNs on a target is unsupported. First determine the intended topology from camcontrol devlist -v; never guess a target number based on a device name. Confirm that the hardware was physically connected or enabled and that the controller supports the operation. Capture the before state, perform the narrowest rescan during an approved window, then capture the device list and kernel messages again.
A rescan is not a repair for a consistently failing disk. Repeated rescans can create noisy evidence, stress an unstable transport, and complicate incident timelines. A device may be discovered but still not be usable: it may lack a peripheral driver, have no media, report an unsupported logical block size, or have no GEOM provider because a later layer did not attach. Follow the attach logs to find the first failed transition.
NVMe illustrates why transport identity matters. The FreeBSD nvme(4) and nda(4) manuals describe NVMe namespaces, where a namespace is roughly analogous to a SCSI LUN. When using nda, the namespace ID is mapped to the CAM LUN and can be shown by camcontrol devlist. When the nvd driver is selected instead, the exposed block-device path differs. Inspect hw.nvme.use_nvd, loaded drivers, and the actual device name before using CAM-specific commands; do not infer the path only from the fact that the media is NVMe.
Protect consumers and data
Before detaching or replacing anything, inventory GEOM consumers. Check mount, swapinfo, zpool status, gpart show, and GEOM topology. A disk can host a mounted filesystem, a swap provider, a pool member, a mirror, a VM disk, or an iSCSI backend. EBUSY is a protective refusal when a provider is still open; forcing teardown can destroy in-flight writes or metadata. Follow the normal unmount, swap-off, pool-export, and path-management procedures appropriate to the consumer, and verify each layer has released the provider.
The device-unit number is not a durable asset identifier. Preserve hardware identity, controller path, serial or WWN, firmware, and the mapping between kernel names and labels. Stable GEOM labels can help configuration survive discovery-order changes, but labels do not replace an inventory of the physical device and its current path. When a replacement is performed, confirm the new device’s identity and capacity before allowing a mirror resilver, pool import, partitioning, or filesystem creation.
A recovery decision with acceptance criteria
Classify the incident before acting: discovery failure, transport or controller error, device readiness/media error, GEOM/provider issue, or filesystem/application failure. State what evidence supports that classification and what would disprove it. If data is still readable but the device is degrading, prioritize verified backups and minimizing writes. Do not run a repair utility against unstable media without a copy or recovery plan.
After repair or hardware replacement, verify each boundary: the controller attaches without reset loops; CAM identifies the expected bus and LUN; the peripheral driver exposes the correct disk; GEOM maps the intended provider; partitions and labels resolve; the filesystem or pool reports healthy; and the application completes a representative read/write test. Monitor the next reboot and a sustained workload, because a device that attaches once may still have a power, cable, firmware, or thermal fault.
CAM diagnostics are most valuable when they preserve hierarchy. camcontrol tells you about CAM-visible devices and commands; GEOM tools reveal block-provider relationships; filesystem and application tools establish whether service recovered. Keeping those findings separate turns a generic “disk error” into a testable diagnosis and reduces the risk that recovery work destroys the data it was meant to protect.
Related:
- Understanding GEOM: FreeBSD’s Modular Storage Framework
- Fixing ‘Device Busy’ on a FreeBSD GEOM Provider Without Forcing Data Loss
Sources: