Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux NVMe Multipath: Path Selection, ANA State, and Failover Evidence

Diagnose Linux NVMe multipath with subsystem and path inventories, ANA state, I/O policy, failover tests, and application-level acceptance evidence.

NVMe multipath lets Linux represent multiple controller paths to the same namespace as a shared storage resource. A host may reach one namespace through several controllers, fabrics, or paths, improving availability and sometimes distributing I/O. It does not create redundancy by itself: the paths must reach the same intended namespace, the storage target must report path state correctly, and the application must tolerate the latency and error behavior of path loss.

The first troubleshooting task is to distinguish a namespace from its paths. Seeing multiple /dev/nvme* devices does not prove that they are duplicate paths to the same logical storage. Conversely, a single multipath-facing device can hide several controller paths. Inventory subsystem identity, namespace identity, controller relationships, ANA state, and the active I/O policy before changing anything.

Confirm multipath support and map devices

Use the NVMe command-line tools and sysfs to inspect the running host. Tool output varies with nvme-cli version and kernel support, so save the full output with timestamps:

uname -r
nvme version
nvme list
nvme list-subsys
nvme list -v

nvme list-subsys is particularly useful for showing subsystem membership and controller paths. Correlate namespace identifiers, controller names, transport addresses, and path state. Do not use device node numbering as a durable identity; names such as nvme0n1 can change after reboot or hotplug. Use stable identifiers and the storage fabric’s own mapping when correlating a Linux device with an array volume.

On systems that expose the kernel parameter, inspect whether multipath is enabled and which path policy is selected:

cat /sys/module/nvme_core/parameters/multipath 2>/dev/null
for f in /sys/class/nvme-subsystem/nvme-subsys*/iopolicy; do
  printf '%s: ' "$f"
  cat "$f"
done

Missing sysfs files may indicate a kernel/configuration difference, not necessarily a failure. Consult the running kernel documentation and distribution configuration. The upstream implementation provides policies such as NUMA-aware selection, round-robin, and queue-depth-based selection; exact availability and defaults should be confirmed for the deployed kernel. A policy is not a failover test and cannot repair a physically broken fabric.

Understand ANA and path state

For NVMe over fabrics, Asymmetric Namespace Access (ANA) communicates access state for paths to a namespace. Optimized and non-optimized states influence path preference; inaccessible or persistent-loss states indicate that a path should not be treated as a healthy equivalent. Multipath selection follows the namespace’s reported availability and selected policy. Do not assume every path has equal latency or equal access characteristics.

Inspect the subsystem listing and controller state before interpreting a path as failed. A path may be connected but non-optimized; a controller may exist while its namespace path is unavailable; transport and target events may precede a state transition. Correlate kernel journal messages, fabric counters, switch/transport telemetry, and array-side path reports. A single nvme list snapshot is not enough to establish when a transition occurred.

Path policy changes are tradeoffs. NUMA-aware selection can favor a path local to the submitting CPU; round-robin can distribute selections; queue-depth policy can prefer a less busy path. Real throughput and tail latency depend on target behavior, path topology, queueing, device firmware, and workload size. Measure with representative I/O and monitor per-path counters where supported before changing policy.

Diagnose a degraded or missing path

Begin with read-only inventory and recent logs:

nvme list-subsys
nvme list -v
sudo journalctl -k --since '-30 minutes' | grep -Ei 'nvme|ana|I/O error|reset|timeout'
ls -l /dev/disk/by-id/ | grep -i nvme

The journal filter is a convenience, not a complete diagnostic. Preserve unfiltered kernel messages around the event because transport names and failure wording vary. Check whether the namespace is still available through another path and whether the application sees its expected multipath device. Avoid starting repair, format, or namespace management commands during a path incident unless a storage administrator has confirmed the target and action.

Compare host observations with the fabric and target. Verify physical link, zoning or discovery configuration, authentication, controller health, ANA state, and whether a maintenance event was in progress. A host can report a path failure while the target has a separate subsystem incident; conversely, a target may believe a path is usable while a switch drops traffic. Establish which layer first recorded the transition.

For a path that reappears, validate stable namespace identity before allowing workloads to resume. A newly discovered namespace with a similar model string is not proof that it is the original volume. Confirm identifiers from the storage system and the Linux subsystem view. Multipath is designed to merge paths to the same namespace, not to merge unrelated storage merely because their capacities match.

Do not treat path count as a direct availability percentage. Several controller paths can still share a fabric switch, power domain, target port, or storage controller, leaving a common failure point. Map each Linux path back to independent physical and logical components with the storage team. A topology diagram that records host initiator, transport endpoint, switch, target controller, and namespace identifier makes “redundant” testable. Validate that the intended failure domain is actually separated before claiming a path pair protects against it.

Likewise, a successful path-policy change should not be judged only by aggregate throughput. Compare latency distribution, retries, queue depth, and application-visible errors under the same workload and path topology. A policy that improves average throughput can worsen tail latency or concentrate traffic on a path with a slower target. Keep the original policy and inventory so rollback restores the known configuration rather than an assumed kernel default.

Test failover without risking production data

A proper failover exercise starts with a non-production namespace or an approved maintenance plan. Record baseline path state, I/O latency, error counters, application health, and ongoing writes. Generate bounded, checksummed I/O to a scratch workload while removing one path through the supported fabric or target maintenance procedure. Never remove a live path by blindly unbinding controllers on a production system.

Observe whether the remaining path stays usable, how long I/O stalls, whether the application retries safely, and whether the failed path returns to service. Capture nvme list-subsys, kernel logs, and workload-level latency around each state transition. Test recovery in the reverse direction as well: restoring a path should not cause duplicate identities, persistent errors, or unexpected path flapping.

Multipath does not provide exactly-once application operations. An I/O that times out at the host may have reached the target even if the completion was lost. Filesystem, database, and application recovery semantics remain important. Use checksums or application-level validation on disposable data, and verify that the application correctly handles latency spikes and retriable errors.

Operational acceptance and change control

Document the expected number of paths for each namespace, their transports and ANA states, the selected I/O policy, and the owner of the storage mapping. Define thresholds for path degradation and alert on path-count changes, inaccessible states, resets, and prolonged queueing. Distinguish a planned non-optimized path from a missing path so routine target maintenance does not become either a false emergency or a silent availability loss.

Before a kernel, firmware, HBA, or fabric change, retain a baseline inventory and plan a canary. Afterward, compare subsystem membership and run an approved I/O validation. The acceptance check should show the intended namespace identity, all expected paths, the expected policy, successful I/O across a controlled single-path loss, bounded application impact, and clean restoration. Keep a rollback path that does not require writing to or recreating the namespace.

NVMe multipath is a coordinated host, transport, and target feature. Correct diagnosis depends on identity and state, not merely device count. Inventory every path, interpret ANA and policy together, test failure under controlled conditions, and evaluate the result from the application’s perspective as well as the kernel’s.

Related:

Sources:

Comments