Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux SCSI Error Handling: Timeouts, Resets, and Device Recovery

Interpret Linux SCSI timeout and error-handler activity across command retries, device and bus resets, host recovery, and application-visible failures.

The Linux SCSI midlayer coordinates commands between block or subsystem clients and host-bus drivers. When a command times out or fails, the error-handling path can abort a command, reset a device or bus, escalate to host recovery, and retry or fail requests. The log message is a protocol and recovery event, not proof that a disk has failed.

A timeout can originate in media, firmware, a transport, a host adapter, a target, power management, or an overloaded queue. A successful reset can restore command processing while still losing in-flight application operations. Diagnose the command, target, host, retries, reset scope, and service impact before changing timeout values.

Map host, channel, target, and LUN

A SCSI address uses host, channel, target, and logical-unit components. The host identifies the adapter path, while target and LUN identify a device behind it. Multipath, USB storage, SAS expanders, virtual disks, and device-mapper can make the physical relationship more complex than one /dev/sdX name.

lsscsi -g
ls -l /sys/class/scsi_device/
journalctl -k -b --no-pager
lsblk -o NAME,HCTL,TYPE,SIZE,MODEL,SERIAL

Utilities and columns vary. Correlate the H:C:T:L path with the sysfs device chain, host driver, transport, block device, and multipath map. A reset at the host level can affect several targets; replacing the disk named in the last log line may not address a failing HBA or cable.

Preserve the exact command opcode, target status, host status, driver status, sense data, and timeout message when the kernel provides them. Sense data can distinguish not-ready, medium, hardware, and unit-attention conditions. A reset can cause expected unit attention, so separate recovery noise from repeated underlying failures.

Why commands enter error handling

Each SCSI command has a timeout and can complete with a transport or target status. If it does not complete within the expected interval, the midlayer invokes error handling. The SCSI error-handler thread coordinates recovery outside the ordinary command completion path so that the host driver can quiesce or reset hardware safely.

Recovery levels can include command abort, device reset, bus reset, and host reset, depending on driver support and the failure. Escalation broadens the disruption. A device reset may interrupt one LUN; a bus or host reset can disturb unrelated devices and trigger a burst of retries. The exact callback behavior is driver-specific and should be checked against the host driver documentation.

Timeout duration is not a universal performance knob. Increasing it can prevent premature recovery for a slow device, but it can also extend application stalls and delay failure detection. Decreasing it may create reset storms during legitimate long operations such as firmware updates, error recovery, or overloaded paths. Use evidence from command latency and device specification.

Retries and application semantics

The kernel can retry some commands after a transient condition or reset. A retry is safe only if the command and higher-level operation preserve their semantics. A read may be repeatable; a write or vendor command can have side effects. Applications generally see block or subsystem completion errors rather than raw SCSI callbacks, so a failed request can leave uncertainty about whether a device performed the operation.

Databases and filesystems rely on ordering and durability protocols above the transport. A reset during a write does not prove that no sectors changed. Follow the filesystem and application recovery path, and inspect logs for journal replay or rejected writes. Do not power-cycle a system or run a repair utility before preserving device-health and kernel evidence unless service restoration requires it.

For tape, optical, scanner, and other non-block SCSI devices, command semantics differ from disks. The same error handler can serve many device types, but application recovery must use the subsystem’s contract. Avoid treating every SCSI timeout as a block-device I/O failure.

Observe recovery and avoid blind tuning

Read sysfs and logs before changing anything. Some kernels expose per-device command timeout or host error-handler deadline attributes; availability and write permissions vary. Record the current value and the driver documentation before considering a change. Do not write a new timeout to a wildcard path across every host.

Correlate error-handler events with device temperature, link resets, SAS expander events, USB disconnects, NVMe/SCSI bridge errors, HBA firmware, and controller queue depth. If errors align with host resets and multiple devices, investigate the shared path. If only one LUN repeatedly returns medium errors while other devices remain healthy, media is a stronger possibility.

Use vendor tools or SMART/SCSI diagnostic pages to collect supported health data, but do not launch destructive tests during production. Monitor retry counts and latency distribution over a representative workload. A low steady-state error count does not prove recovery is correct; test failure paths on a lab system or virtual SCSI target.

Recovery operations and risk boundaries

Rescanning a SCSI host, deleting a device from sysfs, or issuing a reset can remove live block devices and interrupt users. Never copy a rescan or reset command into a production shell without confirming host, target, mounts, multipath state, and vendor procedure. For hotplug storage, use the system’s supported orchestration and verify all paths before removing one.

If multipath is present, distinguish a failed path from a failed logical device. The device-mapper multipath layer can fail over while the SCSI EH path handles one transport route, but recovery timeouts can still stall the map. Capture both path state and kernel recovery sequence. Do not issue a low-level reset that bypasses the multipath manager.

Before firmware or driver changes, ensure the host has a maintenance path, backup, and rollback. A host-adapter reset can disconnect root storage on some systems. Update one layer at a time and record whether timeouts, resets, and application errors changed.

Acceptance criteria

A storage runbook should identify H:C:T:L, host driver, transport path, block or subsystem mapping, command status and sense data, recovery level, retry outcome, and application impact. It should differentiate single-device faults from shared-path failures and define when to escalate to the controller or device vendor.

Test timeout, abort, reset, path failover, and recovery with disposable devices or a virtual target. Confirm that logs identify the failed layer, applications receive bounded errors, and recovery does not leave stale device state. A system is not healthy merely because repeated resets eventually restore I/O.

SCSI error handling is a layered recovery state machine. The useful question is not “should the timeout be longer?” but “which command stopped completing, which recovery level ran, what else shared that path, and what did the application observe?”

Related:

Sources:

Comments