Cluster API MachineDeployment Rollouts: Surge, Availability, and Node Replacement
Plan Cluster API worker-machine rollouts with template rotation, MachineSet ownership, maxSurge and maxUnavailable math, workload drain, and recovery checks.
Cluster API MachineDeployment manages a fleet of worker Machines by reconciling MachineSets and Machine objects. It plays a similar orchestration role to a Kubernetes Deployment, but each Machine represents host infrastructure and a corresponding Kubernetes Node. Updating the machine template can replace nodes gradually, subject to rollout strategy, cloud capacity, bootstrap readiness, node draining, and the behavior of workloads scheduled on those nodes.
A machine rollout is not an in-place update of an operating system image in the general case. Cluster API treats machines as immutable for most spec changes: a new desired template creates a new MachineSet, and old Machines are replaced with new ones. Some infrastructure providers may support selected in-place operations, but that does not make every template mutation an in-place change or guarantee the MachineDeployment will roll out for fields that the provider modifies directly.
Understand the ownership chain
A MachineDeployment owns or adopts matching MachineSets; those MachineSets maintain Machines. The MachineDeployment controller creates a new MachineSet when rollout-causing template fields change, scales the new set up, and scales the old set down according to the chosen strategy. MachineSets and Machines are implementation resources in this lifecycle; editing them manually can conflict with the parent controller’s desired state.
The bootstrap template and infrastructure machine template are references. If a template is immutable or changes are not propagated in-place, create a new template object and update the reference name to trigger the intended rollout. Editing a referenced template in place may not start a replacement rollout. Review the provider’s API and template mutability rules, and update related Kubernetes version and machine image fields in one planned transaction where the provider requires them to match.
The Cluster API version and infrastructure provider shape are release-dependent. Before applying a manifest or changing a controller, inspect the installed CRDs and provider documentation. Do not copy an API version from a different management cluster or assume a provider-specific template field has the same rollout behavior in another provider.
Set availability and surge intentionally
For a rolling update, maxUnavailable sets the maximum number of Machines that may be unavailable during the update; maxSurge sets the maximum number of additional Machines above the desired replica count that may be scheduled. Values can be integers or percentages. CAPI documents percentage rounding behavior: unavailable counts round down and surge counts round up. The two settings cannot both be zero. Their defaults and API validation must be confirmed against the CAPI release in use.
For a small fleet, percentages can produce unintuitive numbers. With three desired Machines, maxUnavailable: 30% rounds down to zero, while maxSurge: 30% rounds up to one. With a large fleet, that same percentage can allow a substantial replacement wave. Choose values by modeling actual machine count and a failure scenario, not by copying a percentage from a sample.
Surge requires spare quota and capacity in the target infrastructure. The new Machine may need an IP address, subnet slot, instance quota, disk allocation, and bootstrap resources before an old Machine can be removed. If the region has no capacity, the rollout can stall while preserving availability. maxUnavailable can allow progress without surge capacity but creates a deliberate temporary capacity reduction. Neither setting overrides cloud-provider limits or Kubernetes scheduling constraints.
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachineDeployment
metadata:
name: workers
spec:
clusterName: production-cluster
replicas: 6
selector:
matchLabels:
pool: workers
template:
metadata:
labels:
pool: workers
spec:
version: "v1.XX.Y"
bootstrap:
configRef:
apiVersion: bootstrap.cluster.x-k8s.io/v1beta2
kind: KubeadmConfigTemplate
name: workers-bootstrap-v2
infrastructureRef:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: ExampleMachineTemplate
name: workers-infra-v2
rollout:
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
The example uses the current core and Kubeadm bootstrap API versions; the infrastructure API version, template kind, and names are provider-specific placeholders and must match the installed releases. v1.XX.Y is also a placeholder Kubernetes version. Validate the rendered manifest with the management cluster’s schema and inspect the provider’s upgrade guide. The maxSurge and maxUnavailable pair expresses the example’s availability intent, but it does not prove that surge capacity exists.
Account for workload disruption separately
Machine availability and Pod availability are different layers. A MachineDeployment can decide to delete a Machine while Kubernetes workloads on its Node still need to drain. Node drain behavior, pod disruption budgets, termination grace periods, local data, DaemonSets, and provider deletion timeouts affect how quickly a replacement can proceed. A strict PodDisruptionBudget can block eviction even when the MachineDeployment has room to surge.
Before replacing worker nodes, inspect workload replica counts, disruption budgets, topology constraints, anti-affinity, storage attachment limits, and node selectors. A new node can be Ready but unsuitable for a workload because of taints, labels, architecture, or local device availability. A Machine can be provisioned successfully while node bootstrap or CNI installation fails. Define what “ready” means at the infrastructure, bootstrap, Node, and application layers.
Do not lower maxUnavailable to zero and assume there can be no downtime. If the new Machine never becomes Ready, the rollout may stall; if application replicas are already unhealthy or concentrated on the same failure domain, draining one node can still disrupt service. Availability requires adequate replica placement and valid PDB configuration as well as conservative machine rollout parameters.
Roll out a machine template change safely
Start by capturing the MachineDeployment, owned MachineSets, Machine conditions, Node status, and provider resources. Confirm the old and new templates, Kubernetes version, image version, subnet, instance type, and bootstrap configuration. Ensure the management cluster can reach both the old and new provider APIs. Check quotas and capacity in each zone, and determine how maxSurge interacts with autoscaling and budget limits.
Make the template change in a reviewable commit. If the provider requires a new immutable template object, create it first, then update the MachineDeployment’s reference. Observe the first new Machine from provisioning through bootstrap, Node registration, readiness, and workload placement. Verify that the controller creates the expected MachineSet and does not roll out for unrelated metadata changes. Continue in bounded waves and pause when conditions do not match the runbook.
Use Cluster API status and conditions as evidence, but also inspect infrastructure provider events and Kubernetes Node conditions. In the current v1beta2 API, a Machine in the Provisioning phase may have a blocked BootstrapConfigReady, InfrastructureReady, or NodeReady condition; each points to a different owner and likely needs a different diagnostic. Review controller logs and kubectl describe events before deleting a Machine or MachineSet manually. Manual deletion can cause a controller to create replacements, but it can also bypass safeguards or lose useful evidence.
Interpret rollout progress and stalled states
MachineDeployment status includes replica counts and conditions that describe desired and observed state. Compare desired replicas, up-to-date replicas, ready replicas, available replicas, and old MachineSet capacity. A rollout can have new replicas but still fail to progress if Machines are not ready, old Machines cannot drain, a PDB blocks eviction, or provider capacity is unavailable. Inspect the current v1beta2 conditions, their reasons and messages, controller events, and logs to identify the first stalled reconciliation. The deprecated v1beta1 progressDeadlineSeconds field is not part of the current v1beta2 API; if operating an older served API, interpret its progress-deadline signal according to that release rather than assuming it exists in v1beta2.
Common bottlenecks include a capacity shortage in one availability zone, quota limits, a stale bootstrap reference, an incompatible machine image and Kubernetes version, a cloud-init script that never completes, a CNI failure, a PDB with no permitted disruptions, and a storage volume that cannot detach. Identify the first failing boundary rather than increasing maxSurge blindly. Increasing surge can worsen quota pressure; increasing maxUnavailable can convert a safe stall into a capacity incident.
If the rollout is paused, confirm who paused it and why. Cluster API rollout tooling can pause or resume a resource by changing the spec’s paused state. A pause stops reconciliation; it does not roll back already created Machines or stop cloud-side health checks. Resume only after the underlying failure is understood and the desired template is still correct.
Separate rollout from health remediation
MachineHealthCheck is a distinct controller resource that can remediate Machines meeting configured unhealthy conditions. It can delete or replace a Machine owned through a MachineSet, but it is not an alternative rollout strategy and does not fix a bad template. Carefully define remediation thresholds and allowed concurrent remediations so a transient cluster-wide outage does not trigger widespread replacement.
During a rollout, the old and new MachineSets coexist. Health remediation and rollout can interact, so inspect both controllers’ state and avoid duplicating deletion actions. Keep remediation budgets conservative and test behavior in a non-production cluster. A failed upgrade may require stopping rollout progression, repairing the new template or provider, and then choosing whether to resume or revert; blindly allowing health remediation can create extra churn.
Define completion and rollback criteria
A successful infrastructure rollout requires more than all desired Machines existing. Confirm each new Node is Ready, has expected labels and taints, belongs to the intended zone and pool, passes network and storage checks, and can host representative workloads. Verify old Machines are drained and deleted only after replacement capacity is proven. Check application SLOs and disruption metrics throughout the rollout.
Rollback is not always the inverse of changing a template reference. The old machine image may no longer be available, Kubernetes version skew may restrict downgrade, and data or configuration changes may be irreversible. Keep the previous template and image accessible through the rollback window, record any version-skew constraints, and rehearse a replacement with the old configuration. Preserve MachineSet history long enough to support diagnosis without retaining stale infrastructure indefinitely.
MachineDeployment rollouts provide declarative host replacement, but safe operation depends on understanding its full control chain: template identity, MachineSet ownership, provider capacity, bootstrap, node readiness, drain, and application disruption. Model surge and unavailability numerically, observe each layer, and let controllers converge instead of hand-editing their child resources during a rollout.
Related:
- Kubernetes Upgrades: Version Skew, Component Order, and Safe Rollout
- Kubernetes Graceful Node Shutdown: Kubelet Budgets, Priority, and Limits
Sources: