Cluster-Aware Updating: A Windows Server Rolling-Maintenance Runbook
Use Cluster-Aware Updating safely by validating readiness, choosing an update mode, controlling each run, and proving workloads recover after node maintenance.
Cluster-Aware Updating (CAU) coordinates a rolling update run across Windows Server failover-cluster nodes. For each node, it moves clustered roles away, installs selected updates, restarts when required, and returns the node to service. That sequence reduces the need for an operator to manually repeat maintenance on each member, but it does not make an application highly available by itself. Availability still depends on cluster capacity, role failover behavior, storage and network health, client reconnect behavior, and whether the remaining nodes can safely carry the workload.
Use CAU as an orchestrator within a reviewed maintenance plan. It applies updates and tracks its actions; it does not prove that a database, file service, virtual machine, or client transaction is healthy afterward. Separate cluster state from workload state in both monitoring and acceptance criteria.
Understand the two coordination modes
In remote-updating mode, a separate Update Coordinator runs an on-demand Updating Run against the target cluster. The coordinator is not a member of the cluster being updated. This makes the update process observable from an external management host and is useful where the cluster runs Server Core. The coordinator is an operational dependency: if its session is interrupted, reconnect with the CAU tools and inspect the active run and report before deciding whether a documented recovery action is needed. Do not start a second run while the first might still be active.
In self-updating mode, the CAU clustered role runs on the cluster being serviced and a schedule is configured. As the run proceeds, the role can move so the node currently hosting the coordinator can itself be updated. This enables a scheduled end-to-end process, but it also makes the cluster and its CAU role part of the coordination path. Install the Failover Clustering tools on all nodes for self-updating mode. Remote mode requires the tools on the coordinator; the documentation notes that some features are unavailable if the tools are not installed on cluster nodes.
Neither mode means “patch every node regardless of health.” The run uses a plug-in to identify and apply updates. The built-in Microsoft.WindowsUpdatePlugin works through Windows Update Agent and can use Windows Update, Microsoft Update, or an on-premises WSUS source. CAU’s default behavior selects important general distribution release updates; optional plug-in parameters can change the query. Additional plug-ins, including the hotfix plug-in, have their own requirements. CAU is therefore not a substitute for deciding which updates are approved, testing them, or ensuring that all nodes use the same source.
Establish a preflight gate
Before the first run, test cluster updating readiness using the CAU window or Test-CauSetup from a computer with the Failover Clustering tools. Re-run readiness tests after node, hardware, update-source, or relevant configuration changes. Review each reported issue rather than treating the final test result as a blanket approval.
The operational preflight should include:
- All expected cluster nodes are online and cluster health is acceptable. CAU needs sufficient nodes online to retain quorum.
- The cluster name resolves, nodes can communicate with their intended update source, and all nodes use the same update source.
- The coordinator can reach cluster nodes using the documented remote-management configuration.
- If an update may require a restart, firewall policy allows the required remote restart behavior. Group Policy can prevent CAU from enabling the relevant firewall rule automatically.
- The remaining nodes have capacity and compatible placement for the roles being drained. Test actual role movement, not only CPU or memory totals.
- Application owners have approved the maintenance window and know how to validate recovery.
- Any pre-update or post-update scripts are accessible from each node, tested independently, and included in the run timeout budget.
Avoid running another automatic update mechanism against the same nodes at the same time. Microsoft warns that automatically updating individual nodes on a fixed schedule can produce unpredictable results, interruptions, and unplanned downtime when combined with CAU. Use one controlled orchestration path and a clearly assigned owner.
Use a bounded, reviewable Updating Run
Updating Run profiles store reusable settings such as time limits, retry behavior, failed-node limits, and script paths. They do not store cluster-specific credentials; a self-updating profile also does not store the schedule. If you change the profile after configuring self-updating, configure the cluster’s self-updating options again so that the new values take effect. Keep the profile with the change record and review it as executable operational policy, not as a harmless preferences file.
The cmdlet example below follows Microsoft’s Invoke-CauRun parameter model but deliberately does not use -Force, which suppresses confirmation prompts. Replace the cluster name only after validating the target and profile. A full run installs updates; do not use it as a casual discovery command on production.
$cluster = 'CONTOSO-FC1'
$parameters = @{
ClusterName = $cluster
CauPluginName = 'Microsoft.WindowsUpdatePlugin'
MaxFailedNodes = 0
MaxRetriesPerNode = 3
RequireAllNodesOnline = $true
}
Test-CauSetup -ClusterName $cluster
Get-CauPlugin
Invoke-CauRun @parameters
The zero failed-node threshold above is an example policy, not a universal default. It instructs the run to stop if a node failure reaches the configured limit; choose the value from the workload’s supported failure tolerance and cluster topology. RequireAllNodesOnline is useful when the change approval assumes every node starts healthy. It does not replace capacity checks or prove all clustered roles can move. For a production change, preview applicable updates in the CAU UI, record the approved update set, then start the full run in the approved window.
For self-updating mode, add and configure the CAU clustered role through the CAU interface or Add-CauClusterRole, and define the schedule intentionally. Avoid setting a schedule before verifying the run profile and testing how the cluster behaves when one node or the update source is unavailable. A scheduled process still needs alerts, an owner, and a documented path to inspect failures.
Observe each node transition
CAU performs a scan, moves roles off a node, applies updates and dependencies, restarts if necessary, returns the node to service, and can move roles back. It also checks quorum maintenance, looks for additional updates that become applicable after the first set, and saves a report. Monitor the run from the coordinator or query the cluster with Get-CauRun for a summary of an active self-updating run.
Use a time-bounded run profile. Microsoft’s options include StopAfter, WarnAfter, MaxRetriesPerNode, MaxFailedNodes, RequireAllNodesOnline, and RebootTimeoutMinutes. Their defaults can vary by option; inspect the profile and the documentation for the Windows Server version in use rather than relying on memory. The overall timeout includes pre-update scripts, update installation, restarts, and post-update scripts. A reboot timeout that is shorter than the actual firmware or service startup path can mark a node failed even if it later becomes healthy.
Do not use forced drain parameters as routine “make it continue” switches. CAU documents that a forced drain can move roles even if a group cannot move because no other node can host it or the group is locked. If ordinary drain fails, that is important evidence: inspect placement constraints, role state, quorum, storage, and workload dependencies. A force that makes the run advance can also take the workload offline.
For every node, correlate CAU report entries with Failover Clustering events, Windows Update history, reboot completion, and application health. If a node fails or the run stops, pause new maintenance and determine whether the cluster is in a supported state before resuming or recovering the run. Do not infer that a successful CAU task means the update is correct for the application.
Failure patterns and recovery boundaries
- A role cannot move. Check possible owners, placement rules, node capacity, locked groups, storage paths, and application-specific dependencies. Do not force-drain simply to finish patching.
- Updates are not uniform. Compare update sources, approval state, proxy configuration, and plug-in query arguments across all nodes. An update that is visible on one node but not another may reflect source configuration, not CAU randomness.
- Restart never completes inside the run window. Inspect node boot, service startup, and management reachability. Increase a timeout only after understanding the actual recovery path; a larger number does not make a broken restart healthy.
- The coordinator disappears. Reconnect to the cluster and review the CAU run state and report before restarting commands. Avoid starting a second run while the first may still be active.
- The cluster remains online but the service is degraded. Validate the workload, storage, replication, client connectivity, and application-level health separately. Cluster quorum is not an application health check.
- The run repeatedly fails on one node. Treat this as a node-specific investigation. Compare installed updates, pending reboot state, event logs, update-source reachability, and CAU trace/report details before lowering failure thresholds.
Define measurable acceptance criteria
Before the window, record cluster membership, quorum state, role owners, workload capacity, approved update scope, and the profile version. During the run, record which node is being drained, whether roles moved, update/reboot status, the time it took for the node to rejoin, and any retry or warning. Afterward, require every intended node to be online, the run report to be reviewed, expected updates to be present, and every critical clustered workload to pass its own health check.
A useful test in a non-production cluster includes a normal run, a node that has a deliberately unavailable update source, and a workload that cannot move because of placement or capacity constraints. The purpose is to learn how CAU reports failure and how operators recover, not to test forced draining on a production service. Measure client-visible interruption, not only time spent in maintenance mode. If the application has a recovery point or replication requirement, validate it through the workload’s own tools.
CAU makes repeatable rolling maintenance possible. Production safety comes from a known update source, a reviewed run profile, enough healthy capacity, transparent failure handling, and an acceptance check that proves the service recovered. A green CAU summary is necessary operational evidence, not the whole service-level verdict.
Related:
- Windows Server Failover Clustering: Quorum, Witnesses, and Safe Maintenance
- Setting Up WSUS for Centralized Windows Update Management
Sources: