Skip to content
WindowsDeep Dive Published Updated 9 min readViews unavailable

Windows SMB Operations: Multichannel, Durable Opens, and Transparent Failover

Engineer resilient Windows file services with SMB Multichannel, RDMA, durable reconnects, continuously available shares, and measurable failover tests.

An SMB share can be reachable while still being operationally fragile. A single network path can limit throughput and turn one cable failure into an application error. A client that opens a file on a clustered file server can also lose that open during node maintenance unless the share, protocol, server topology, and client support transparent reconnection. Windows SMB 3 includes capabilities for both problems, but they solve different layers and have different requirements.

This guide treats SMB as a client-server protocol, not as a mapped-drive feature. SMB Multichannel provides multiple network connections for a session when the endpoint capabilities and network layout permit it. Durable opens let a client attempt to reconnect an open after a temporary disconnect. Continuously Available (CA) shares on a supported failover cluster enable SMB Transparent Failover across node maintenance or failure. None of these features replaces application-level transaction recovery, backups, or a test that proves the workload survives.

Start with the negotiated SMB connection

Both the SMB client and SMB server participate in capability negotiation. A Windows computer can act as either role, and feature availability depends on the SMB dialect negotiated for that connection, the Windows versions, adapter features, share configuration, and file-server topology. Do not infer protocol features from the operating system edition alone. Check a real session from the client and the server that owns the share.

# Run on a client while the share is actively in use.
Get-SmbConnection |
  Format-Table ServerName, ShareName, Dialect, NumOpens

# Show selected and eligible client/server interface pairs.
Get-SmbMultichannelConnection -IncludeNotSelected |
  Format-List *

Get-SmbConnection reports the dialect for an active client connection. Get-SmbMultichannelConnection shows channel selection and interface capability information; it is more useful during an ongoing transfer than when no share session exists. On the server, inspect sessions and interfaces as well, because a healthy client adapter list cannot prove that the server advertised the expected interfaces or that the network allows the paths.

Get-SmbServerNetworkInterface
Get-SmbServerConfiguration | Select-Object EnableMultiChannel
Get-NetAdapter | Format-Table Name, Status, LinkSpeed, InterfaceDescription
Get-NetAdapterRss
Get-NetAdapterRdma

Run commands in an elevated PowerShell session where required, and note which host each command targets. In a managed environment, use the appropriate remote PowerShell session or management endpoint rather than accidentally collecting data from the workstation running the shell.

SMB Multichannel: bandwidth and path resilience

SMB Multichannel is part of SMB 3 and is enabled by default on supported Windows SMB clients and servers. It discovers multiple usable network paths and can establish more than one connection for a single SMB session. Depending on adapters and their capabilities, it can aggregate bandwidth, use multiple connections over RSS-capable interfaces, or use RDMA-capable adapters with SMB Direct. It can also let a transfer continue when one of several active network connections is lost.

More paths do not automatically mean higher application throughput. If storage is the bottleneck, adding channels will not make writes commit faster. A single adapter may expose multiple receive queues through Receive Side Scaling, while multiple adapters may have different subnets, speeds, MTUs, or routing. SMB selects channels based on discovered capabilities and connectivity, not on an operator’s diagram. Verify the actual selected paths, counters, link health, and workload behavior.

Avoid configuring network changes solely to maximize the number of channels. Incorrect RSS or RDMA configuration, asymmetric routing, firewall filtering, driver defects, and mismatched network policies can reduce reliability or cause SMB negotiation to fail. Microsoft documents NIC teaming as one supported topology and notes that teaming retains its own failover capability; the correct design depends on the operating system version and adapter stack. Validate the specific vendor-supported design before combining teaming and Multichannel.

When a transfer is active, compare selected channels with adapter and SMB performance data:

Get-SmbConnection | Select-Object ServerName, ShareName, Dialect
Get-SmbMultichannelConnection
Get-NetAdapterStatistics

Treat this as evidence about a particular connection at a particular time. A second test should disconnect one planned path in a lab while a representative workload is running, confirm the other channel remains selected, and check that the application reports no failed or partial operation. Do not disconnect a shared production interface without proving which other workloads and management paths depend on it.

SMB Direct is a distinct RDMA path

SMB Direct uses network adapters that support Remote Direct Memory Access. On supported Windows Server systems, it is automatically configured and enabled by default. RDMA can provide high throughput with low latency and lower CPU overhead for suitable server workloads, but it requires compatible adapters, drivers, switches, network configuration, and client/server support. An adapter’s marketing specification is not proof that an SMB Direct connection is established.

Check Get-NetAdapterRdma, Get-SmbClientNetworkInterface, and Get-SmbServerNetworkInterface on the relevant endpoints. During a controlled transfer, inspect Get-SmbMultichannelConnection and SMB Direct performance counters. A connection can use ordinary TCP channels even when an RDMA NIC is present if its capability is not discovered or the RDMA path is not usable. Do not disable SMB Direct as a generic troubleshooting shortcut; isolate the path in a lab and compare repeatable workload measurements.

Use a dataset larger than the available cache when benchmarking, keep test conditions comparable, and include the storage subsystem in the measurement. A small repeated copy can mostly measure memory cache rather than sustained network or disk behavior. Record bytes per second, CPU, storage latency, active channels, and retransmission or interface error evidence. Report the exact client, server, SMB dialect, NIC driver, and network design with the result.

Durable reconnect is not continuous availability

The SMB protocol can preserve some file opens for reconnection after a network disconnect. A durable open is a server-side handle that may be re-established within a bounded timeout if the client can reconnect and the server still has the open. This helps with transient transport interruptions; it is not a guarantee that any open will survive any outage. The server can close a preserved handle after its timeout or under administrative action or resource pressure, and an application still needs to handle a failed operation.

SMB 3 also defines persistent handles for continuously available clustered shares. In that case, the server cluster can preserve the state needed for a supported SMB client to reconnect to another cluster node after failover. The terms are related but not interchangeable: “durable” describes reconnectable opens, while “persistent” adds clustered failover behavior for CA shares. A normal standalone share, a CA flag by itself, or a disconnected client does not establish that the entire chain is configured for transparent failover.

For a clustered file server, verify all of the following before claiming transparent failover:

  • The SMB server is a supported Windows Server failover cluster with at least two configured nodes and passes cluster validation.
  • The share is configured as Continuously Available. This is normally the default for supported clustered shares, but inspect the actual share.
  • The client uses a supported SMB 3-capable Windows client or server version; older clients may connect but do not get transparent failover.
  • Storage and clustered roles follow the supported architecture. SMB Scale-Out uses Cluster Shared Volume paths; other clustered file-server designs may have different storage access patterns.
  • The application can retry or recover at its own layer and does not assume a successful transport reconnect means a transaction committed exactly once.

Inspect the share configuration on the server:

Get-SmbShare -Name "AppData" |
  Format-List Name, Path, ScopeName, ContinuouslyAvailable, ShareState

A CA share improves continuity for compatible sessions; it does not make an application stateless. A client can reconnect yet still need to retry a request whose completion was ambiguous when the node failed. Applications that write structured data should have a documented recovery path and integrity checks. Confirm that the workload vendor supports the SMB topology and CA behavior before moving databases, virtual machines, or other latency-sensitive state onto it.

Test the system, not the checkbox

Use a non-production cluster and a representative client/application. Start with a baseline transfer and capture the negotiated dialect, CA share setting, channels, server node, storage latency, and application-level results. Then test one network-path loss for Multichannel and a planned clustered role move for Transparent Failover separately. These tests isolate two different mechanisms. Finally, test an unplanned node loss only where the cluster and workload owners have approved the disruption.

During a failover test, measure interruption duration and record whether the application observed an I/O error, paused, retried, or required operator action. Compare the client session before and after: did the server node change, did the client reconnect, and did the share remain continuously available? Validate the file or dataset after the test with an application-aware consistency check. A copy command returning success is useful evidence, but it does not validate database transaction semantics or workload correctness.

For repeatable bulk-copy tests, use a fixed dataset, pre-create the destination path, and capture robocopy’s exit summary rather than judging the console animation. For example:

robocopy C:\SmbTest \\FSCLUSTER\AppData\SmbTest /E /COPY:DAT /R:2 /W:2 /LOG:C:\Temp\smb-test.log

Robocopy exit codes have their own success and difference meanings, so interpret the documented return code rather than assuming every nonzero result is a fatal transfer failure. Compare hashes or application-level content checks where appropriate, and ensure the test does not overwrite production data. If using multi-threaded copying, measure a workload representative of many small files separately from large sequential files; thread count can change client CPU and storage pressure.

Troubleshoot in protocol order

When performance or resilience is missing, follow the connection path instead of toggling features blindly. Confirm name resolution and routing, then verify both endpoints expose the expected SMB interfaces. Check that SMB Multichannel is enabled on both sides and that the session negotiates a dialect that supports it. Inspect selected versus unselected channels and RDMA capability flags. Capture SMB negotiation and NETWORK_INTERFACE_INFO exchanges if the interface list is not being shared. Review SMBClient and SMBServer event logs and network adapter counters around the failure timestamp.

If channels are healthy but throughput is low, measure disk latency, CPU, network utilization, and file-size distribution. SMB performance can be bounded by storage, protocol signing/encryption overhead, endpoint filters, or application I/O pattern. Do not turn off security controls as a first step; use controlled performance counters and supported troubleshooting guidance to isolate a cause. If transparent failover fails, check client dialect, share CA state, cluster resource ownership, cluster validation, and whether the outage exceeded the relevant reconnect or application timeout.

Keep a runbook with share names, server/client versions, expected dialect, NIC and RDMA design, CA state, test results, recovery owner, and rollback path. Re-run the acceptance test after SMB dialect policy, Windows updates, NIC driver changes, cluster upgrades, storage changes, or application upgrades. Resilience is the observed behavior of the whole client-to-share path, not an SMB feature bit in isolation.

Related:

Sources:

Comments