Windows Server iSCSI and MPIO: Diagnose Sessions, Paths, and LUN Failover
Trace Windows Server iSCSI from target portal to disk, verify MPIO claiming and path diversity, and investigate drops without risking LUN data.
An iSCSI disk is delivered through a chain of network portals, initiator sessions, target connections, device discovery, multipath ownership, and Windows disk/volume state. A drive disappearing after a network flap can be caused by a target outage, a single-path design, incorrect MPIO claiming, link/VLAN/MTU mismatch, adapter timing, firmware, or disk policy. Reformatting, rescanning repeatedly, or changing a SAN policy before identifying the failed layer can turn a temporary path loss into a data incident.
Treat an iSCSI disk as a SAN-backed device, not as an ordinary local disk. Before repair, record the server, initiator IQN, target IQN, portal addresses, session IDs, number of connections, persistent-login state, disk unique ID, LUN mapping, MPIO policy, adapter paths, and event timestamps. Confirm whether the volume is standalone, clustered, used by Hyper-V, or owned by a storage application. The correct recovery sequence depends on that owner.
Model portals, sessions, connections, and devices
A target portal is a network endpoint through which the initiator discovers a target. An iSCSI session is a logical relationship between initiator and target; it can have one or more connections. MPIO can aggregate multiple supported physical paths to the same logical unit, but multiple IP addresses alone do not prove path redundancy. If both portal IPs traverse the same NIC, switch, upstream trunk, or array controller, one failure can still remove all access.
Gather session and connection state before reconnecting:
Get-IscsiTargetPortal | Format-List *
Get-IscsiTarget | Format-List NodeAddress, IsConnected, IsPersistent
Get-IscsiSession | Format-List InitiatorNodeAddress, TargetNodeAddress,
IsConnected, IsPersistent, IsMultipathEnabled, NumberOfConnections,
InitiatorPortalAddress, SessionIdentifier
Get-IscsiConnection | Format-List *
Get-Disk | Select-Object Number, FriendlyName, SerialNumber, UniqueId,
BusType, OperationalStatus, IsOffline, IsReadOnly, Size
Store this as an incident snapshot along with time and server identity. Match a session to the expected target IQN and LUN using array inventory or vendor tooling. Disk number is not a stable identity across reboots, discovery order, or path changes. Use disk unique ID, serial, LUN mapping, and storage-array records to avoid applying a change to the wrong device.
Inspect MPIO separately from iSCSI login state. A session can be connected while the device has only one usable path. Review the Microsoft Device Specific Module (MSDSM) claim settings, supported hardware, available paths, and load-balancing policy. The array vendor’s supported DSM and firmware matrix takes precedence over a generic setting copied from another SAN.
Prove path diversity rather than counting connections
Document each expected route from host initiator port through NIC, VLAN, switch fabric, target portal, array controller, and LUN. For every path, establish which adapter and target interface carry it. Two target connections over one physical uplink may increase connection count without providing failure independence. Redundant design should use separate approved initiator adapters and storage network paths consistent with the array’s recommendations.
Check link state, interface counters, MTU, VLAN, route selection, DNS/IP configuration where names are used, and whether storage and client traffic are separated as designed. Jumbo frames require consistent end-to-end configuration; a mismatch can produce stalls or retransmissions that are mistaken for disk faults. Flow control or RDMA-related settings must match the supported fabric design. Do not change offloads, RSS, MTU, or teaming while the host owns production LUNs without a maintenance plan and storage/network owner approval.
Windows can establish multiple connections within a session when the target and initiator support it. MPIO’s job is to handle multiple device paths according to supported hardware and policy. Do not assume that multiple connections in Get-IscsiSession are equivalent to multiple independently managed MPIO paths. Compare the session properties with mpclaim, MPIO cmdlets, Device Manager, and the storage vendor’s path view.
Verify MPIO claiming and device support
MPIO must claim the intended device using a compatible device-specific module. If MPIO is installed after the disk is discovered, the device may not be claimed until the documented rediscovery/update sequence is performed. If the array presents an unsupported hardware identifier or a third-party DSM is installed incorrectly, the host can report duplicate devices, missing paths, or inconsistent load balancing.
Collect read-only MPIO state first:
Get-WindowsFeature -Name Multipath-IO
Get-MPIOSetting | Format-List *
Get-MSDSMAutomaticClaimSettings | Format-Table -AutoSize
Get-MSDSMSupportedHW | Format-Table -AutoSize
Get-MPIOAvailableHW | Format-Table -AutoSize
Get-MSDSMGlobalDefaultLoadBalancePolicy
mpclaim.exe -s -d
Check whether the target’s hardware ID is supported and whether MSDSM or the vendor DSM owns it. Automatic claiming is a configuration decision, not a generic repair. If policy changes are required, use the array vendor’s exact guidance, review the affected LUNs, stage the change, and understand whether a reboot or device rediscovery is required. Do not mix DSMs or add broad hardware claims based only on a similar product name.
Load-balancing policies such as round robin or failover-only have performance and compatibility implications. The policy must be supported by the storage array and its firmware. Validate behavior during a path failure and recovery, including outstanding I/O, not only the healthy-state path count. A high path count does not prove the selected policy distributes traffic as expected.
Persistent login and startup ordering
A target login can be persistent so the initiator attempts to restore the session after reboot. Confirm the exact persistent-target state and favorite portal configuration for every required target. Then verify that the network adapters, VLANs, routes, MPIO driver, iSCSI service, and storage array are ready in a safe order. A dependency or delayed-start adjustment should follow supported storage guidance, not be used to conceal a link or target availability issue.
For a planned new login, use the target IQN and portal address confirmed by the storage administrator. Microsoft documents the Connect-IscsiTarget command with -IsPersistent $true for persistent target connections. Treat it as a state-changing example and do not run it until the LUN presentation, portal path, authentication, MPIO behavior, and intended host are confirmed:
$targetIqn = 'iqn.2001-04.example:array01.lun42'
$portalIp = '10.20.30.40'
# Change operation: confirm the target and approved host before running.
Connect-IscsiTarget -NodeAddress $targetIqn `
-TargetPortalAddress $portalIp `
-TargetPortalPortNumber 3260 `
-IsPersistent $true
If CHAP or mutual CHAP is used, handle credentials through the approved secret mechanism and avoid recording them in scripts or transcript logs. Confirm the target’s authentication requirement and initiator identity on both sides. A login failure can be credential, ACL, portal, route, service, or target state; repeatedly changing passwords is not a substitute for reading the iSCSI event evidence.
Diagnose a dropped path without initializing the disk
Common event evidence includes iSCSI connection and target errors, Storport resets, disk I/O retries, surprise-removal events, and filesystem or cluster messages. Correlate events across System, iSCSI, MPIO, storage driver, and cluster logs if applicable. Record event ID, source, device identifier, session/connection state, path, and time. A Disk 157 surprise removal or Storport timeout indicates a storage-path symptom, not permission to initialize or format the disk.
Use network adapter counters and a bounded Windows network trace when the path fails. Compare both endpoints and the network fabric. A TCP connection test to port 3260 can show reachability to a portal but cannot prove that the target accepts the IQN, that authentication succeeds, or that MPIO has claimed the disk. Check array logs for target resets and controller failover at the same timestamp.
If disks appear RAW, offline, or unexpectedly duplicated, stop writes from applications and capture disk identities before changing state. Do not bring a disk online if it is a cluster-managed or shared LUN without coordinating with Failover Clustering. Do not run Initialize-Disk, Clear-Disk, Set-Disk -IsOffline $false, or filesystem repair as a connectivity workaround. Establish whether the operating system is seeing one logical device with multiple paths or multiple presentations of the same LUN.
After the underlying path recovers, verify session persistence, each expected path, MPIO policy, disk unique IDs, volumes, and cluster/application health. A disk can reappear while a workload remains degraded or while a second path is still absent. Observe I/O retries and latency for a defined period, then save a post-recovery snapshot. If the issue recurs, preserve evidence and escalate with the storage/network vendor rather than continuing reset loops.
Use a controlled maintenance and failover test
Test path failure in a lab or approved window by disabling one approved path at a time, observing I/O and MPIO state, and restoring it before testing another. Verify application behavior, cluster ownership, storage-array path counters, and recovery time. Never remove every path to “see what happens” on a production LUN. Planned testing must account for application write-cache behavior, cluster validation, and array controller failover.
Before firmware, driver, or MPIO changes, record versions and the vendor support matrix. Upgrade NIC, HBA, storage, DSM, and array firmware in a coordinated sequence that preserves a viable path. After a firmware update, compare path identifiers and device claims; an unexpected missing disk may reflect hardware compatibility or a rescan issue rather than lost data.
iSCSI/MPIO incident checklist
- Capture IQNs, portals, session IDs, connections, persistent state, disk IDs, and event times.
- Prove independent physical path diversity from initiator NIC to target/controller.
- Verify MPIO feature, supported hardware, DSM ownership, available paths, and policy.
- Check network counters, MTU/VLAN, firmware, and array logs before reconnecting.
- Preserve disk state; never initialize, format, or force online a shared LUN as a connectivity test.
- Validate recovery and application/cluster health after the path is restored.
iSCSI availability is a property of the complete storage path, not the presence of a successful login. Safe operations keep device identity stable, verify that MPIO owns the intended hardware, test genuine failure independence, and preserve every disk and cluster boundary while investigating.
Related:
- Windows Server Failover Clustering: Quorum, Witnesses, and Safe Maintenance
- Hyper-V Virtual Switch Networking: Uplinks, Host vNICs, and VLANs
Sources: