Windows Server Storage Replica: Designing and Rehearsing Failover
Engineer Storage Replica around sync mode, log capacity, initial sync, health evidence, and a rehearsed role switch instead of assuming replication is failover.
Storage Replica is a Windows Server block-level replication feature for protecting volumes between servers or clusters. Its important operational property is also its most easily misunderstood one: a replicated destination is not automatically a second live file server. In the ordinary server-to-server topology, the source volume accepts writes and the destination volume is kept offline while replication is active. A recovery plan must deliberately change replication direction, expose the recovered data, and restore the client path. Replication transports blocks; it does not decide whether an application is healthy or whether users can reach the new site.
Treat the feature as one component in a disaster-recovery design. Document the replication relationship, expected recovery point, application consistency needs, name/share transition, and the people authorized to declare a site failed. Test that sequence before an outage, not during one.
Choose synchronous or asynchronous replication from the recovery objective
Storage Replica supports synchronous and asynchronous modes. In synchronous mode, Windows Server documents crash-consistent volumes and zero data loss at the file-system level during a failure, provided the topology meets the feature’s latency and bandwidth requirements. That is not the same as a guarantee that every application has committed a business transaction. The filesystem can be crash-consistent while an application has unflushed or otherwise non-transactional state. Database engines and other stateful applications still need their own supported recovery and consistency checks.
Asynchronous mode is intended for longer distances and higher-latency links. It continuously sends changes without waiting for the serialized synchronous acknowledgement from the destination. Consequently, the destination can lag when the source writes faster than the link and destination can process the change stream. A failover may therefore expose a point behind the source. Measure replication lag under representative writes and state the tolerated recovery point objective explicitly. Do not describe asynchronous replication as zero-loss protection.
Neither mode is a backup. Replication also reproduces unwanted writes, deletion, and corruption. Microsoft describes Storage Replica as operating at the partition layer, which also means VSS snapshots on the replicated volume are replicated. Snapshots can aid point-in-time recovery, but they do not replace an independently retained backup with tested restore procedures.
Model the data volume and log as separate capacity plans
Each side needs a data volume and a dedicated log volume. The data volumes must be the same size, and the corresponding disks must use matching sector sizes. The log volume records replication activity; it is not a second copy of the data and must not be shared with unrelated workloads. Microsoft recommends faster storage for the log than for the data, and the server-to-server guide documents an 8 GB default log size while warning that the appropriate size depends on topology-test results and organizational workload requirements.
Do not choose a log size by copying another cluster’s configuration. A write burst, link interruption, or slow destination can consume log capacity faster than a quiet lab workload suggests. Run Test-SRTopology using a production-like write pattern and a meaningful evaluation period. Its report measures connectivity, storage and sector requirements, latency, observed write I/O, and estimated initial synchronization time. An idle test volume does not provide representative performance or log-sizing evidence.
An illustrative validation run looks like this; use real server names and volume assignments, and review the report before creating a partnership:
$report = 'C:\SR-Validation'
New-Item -ItemType Directory -Path $report -Force | Out-Null
Test-SRTopology `
-SourceComputerName 'SR-PRIMARY' `
-SourceVolumeName 'F:' `
-SourceLogVolumeName 'G:' `
-DestinationComputerName 'SR-RECOVERY' `
-DestinationVolumeName 'F:' `
-DestinationLogVolumeName 'G:' `
-DurationInMinutes 30 `
-ResultPath $report
This is a validation example, not a performance guarantee. A thirty-minute test is useful only if the workload, storage, and network resemble the conditions the service will experience. Save the generated HTML report with the change record and resolve warnings instead of treating cmdlet completion as proof of readiness.
Establish the partnership and prove initial synchronization
The server-to-server setup uses named replication groups and one source/destination volume pair. A simplified creation example from Microsoft’s documented workflow is:
New-SRPartnership `
-SourceComputerName 'SR-PRIMARY' `
-SourceRGName 'RG-PROD' `
-SourceVolumeName 'F:' `
-SourceLogVolumeName 'G:' `
-DestinationComputerName 'SR-RECOVERY' `
-DestinationRGName 'RG-DR' `
-DestinationVolumeName 'F:' `
-DestinationLogVolumeName 'G:' `
-LogType Raw
Names and volume paths here are illustrative. Confirm the exact supported configuration for the Windows Server release, edition, storage arrangement, and topology before using a production command. The destination data volume is dismounted by design during active replication. Do not try to make it writable by assigning a drive letter or editing the replicated contents behind Storage Replica.
Initial block copy is a separate operational phase. It may take substantial time, and the partnership is not ready for a controlled failover until the initial synchronization has completed. Inspect partnership and group state on both systems, not merely the fact that the New-SRPartnership call returned an object:
Get-SRPartnership | Format-List *
Get-SRGroup | Select-Object Name, Replicas
Get-WinEvent -ProviderName 'Microsoft-Windows-StorageReplica' -MaxEvents 100 |
Select-Object TimeCreated, Id, LevelDisplayName, Message
Use the documented Storage Replica counters to establish a baseline and trend, including replication transaction counts, maximum log sequence number, and pending or average flush queue length. Event IDs and cmdlet properties provide useful diagnostics, but operational monitoring should also alert on a partnership that stops advancing, an unexpectedly growing queue, failed jobs, and loss of the replication network. Keep timestamps synchronized so source and destination evidence can be correlated.
Make direction changes a controlled failover, not a reflex
Changing direction changes which replication group is the source. Microsoft documents Set-SRPartnership for this operation and specifically warns not to force a direction change before initial synchronization completes. A planned switch should happen only after the recovery side is verified and the application is quiesced according to its own runbook. A disaster switch additionally requires an explicit decision that the former source cannot continue accepting writes; otherwise, operators risk two sites acting on diverging copies.
For a planned role switch, the command shape is:
$switch = @{
NewSourceComputerName = 'SR-RECOVERY'
SourceRGName = 'RG-DR'
DestinationComputerName = 'SR-PRIMARY'
DestinationRGName = 'RG-PROD'
}
Set-SRPartnership @switch
This is not a complete failover script. The current source becomes the destination, and writes are blocked on the originating source as the direction changes. Verify the new partnership state, inspect the Storage Replica event log for the transition and recovery mode, and then follow the documented workload and namespace procedure to make the new source available to clients. For file shares, that may involve shares and DFS Namespaces targets; for other workloads, it may involve application startup, service identity, name resolution, and validation of the recovered dataset. Never promote a second copy merely because the first server is slow to respond.
Failback is another planned direction change, not an automatic undo button. First let replication synchronize in the reverse direction, verify the new source and destination roles, quiesce the workload, and coordinate the name or service transition. If the recovery copy was active for hours, it may contain valid writes that were never present at the original site. The runbook must define whether to preserve those writes, reconcile them, or abandon them before moving service back.
Failure cases to rehearse
- The link fails while the source is healthy. Determine whether writes remain available on the source, how lag is measured, and who decides whether to keep serving or fail over. An asynchronous copy may be behind.
- The destination is not mounted. This is expected during normal replication. It becomes a problem only when the runbook assumes the passive volume can be browsed as if it were an ordinary read-only share.
- The log fills or storage stalls. Check log and data volume health, free space, write rates, and replication events. Do not delete or alter log files manually.
- Initial copy is incomplete. Do not switch direction or promise recovery at the target. Restore connectivity and allow synchronization to finish, then validate the state.
- Failover succeeds but clients fail. Replication does not automatically guarantee that DNS, DFS, shares, application dependencies, credentials, or client caches point to the recovery server.
- The replicated blocks are valid but the application is not. Start the workload through its supported recovery path and require application-level integrity checks before declaring service recovered.
An acceptance drill should measure recovery time from the declared incident to a successful client transaction, not just the time taken by Set-SRPartnership. Record the last confirmed synchronization point, direction, role transition, application checks, client name transition, and the person who authorized each irreversible decision. Restore the original direction only after reverse synchronization and another approved change window.
Related:
- Windows DFS Namespaces and Replication: Referrals, Convergence, and Recovery
- Windows Server Failover Clustering: Quorum, Witnesses, and Safe Maintenance
Sources: