Kubernetes Topology Spread Constraints: Design for Failure Domains
Use maxSkew, eligible domains, minDomains, and hard or soft placement to spread replicas safely across Kubernetes zones and nodes.
Topology spread constraints tell kube-scheduler how to place incoming Pods across labeled failure domains such as zones or hosts. They are a placement policy, not a replica manager: the scheduler does not move running Pods when topology changes, and a balanced initial rollout is not an availability guarantee. Reliable behavior starts with correct node labels, a selector that matches the intended Pods, and a deliberate choice between strict balance and placement flexibility.
Define the failure domains and the matching population
A domain is a distinct value for a node label key: for example, each value of topology.kubernetes.io/zone is a zone, and each value of kubernetes.io/hostname is a host. Prefer the well-known keys where they describe the real infrastructure, but inspect the labels on every eligible node rather than assuming a cloud provider or cluster installer populated them consistently.
The constraint’s labelSelector identifies the Pods counted for skew. Only Pods in the incoming Pod’s namespace are candidates; topology spread has no cross-namespace selector. The selector should match the workload’s Pod-template labels and be narrow enough not to combine unrelated applications in that namespace. A selector that does not match the incoming Pod itself can schedule it without including it in the counted population. The incoming Pod’s affinity and node selector, tolerations, topology labels, and storage requirements all affect where it can actually run.
This Deployment is a runnable scheduling smoke test. The pause image has no application traffic; replace it with your approved image and measured resource requests before production.
apiVersion: apps/v1
kind: Deployment
metadata:
name: topology-demo
spec:
replicas: 6
selector:
matchLabels:
app: topology-demo
template:
metadata:
labels:
app: topology-demo
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
nodeAffinityPolicy: Honor
nodeTaintsPolicy: Honor
labelSelector:
matchLabels:
app: topology-demo
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
nodeAffinityPolicy: Honor
nodeTaintsPolicy: Honor
labelSelector:
matchLabels:
app: topology-demo
containers:
- name: pause
image: registry.k8s.io/pause:3.1
resources:
requests:
cpu: 10m
memory: 16Mi
With three eligible zones and enough capacity, the hard zone constraint keeps the matching Pod counts within a skew of one; six replicas therefore normally distribute two per zone. The host constraint is only a preference: it asks the scheduler to favor nodes that improve host-level spread, but it does not forbid a placement when resources or other scheduling priorities point elsewhere. Multiple constraints are considered together, so a soft preference never cancels a hard constraint.
Choose strict balance or best-effort placement
whenUnsatisfiable: DoNotSchedule makes maxSkew a scheduling guardrail. If placing a Pod would exceed the allowed skew relative to the global minimum, that Pod remains Pending until a feasible domain or capacity becomes available. This is appropriate when the workload must not be concentrated beyond an explicit failure-domain objective, but it can reduce usable capacity during maintenance or an outage.
ScheduleAnyway keeps the constraint soft. The scheduler scores feasible nodes to favor lower skew, alongside other scheduling priorities; the resulting distribution can still be uneven. Use it when keeping useful replicas schedulable matters more than a strict placement ceiling. maxSkew must be a positive integer for either policy.
minDomains is a separate, deliberate fail-closed control. It is valid only with DoNotSchedule and a value greater than zero. If fewer eligible domains exist than minDomains, Kubernetes treats the global minimum as zero when calculating skew. This second manifest requires three eligible zones for all three replicas to run:
apiVersion: apps/v1
kind: Deployment
metadata:
name: strict-three-zones
spec:
replicas: 3
selector:
matchLabels:
app: strict-three-zones
template:
metadata:
labels:
app: strict-three-zones
spec:
topologySpreadConstraints:
- maxSkew: 1
minDomains: 3
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
nodeAffinityPolicy: Honor
nodeTaintsPolicy: Honor
labelSelector:
matchLabels:
app: strict-three-zones
containers:
- name: pause
image: registry.k8s.io/pause:3.1
resources:
requests:
cpu: 10m
memory: 16Mi
If only two eligible zones remain, the zero global minimum plus maxSkew: 1 allows at most one matching Pod in each zone; the third replica stays Pending. That may be correct for a contractual three-zone minimum, but it is often the wrong recovery policy for a service that should continue with reduced redundancy. Before using this field, check API support: before Kubernetes v1.30, minDomains was available only when the MinDomainsInPodTopologySpread feature gate was enabled; that gate is on by default starting in v1.28, but older or customized clusters may have it disabled or may not expose the field.
The examples explicitly use nodeAffinityPolicy: Honor and nodeTaintsPolicy: Honor so the skew calculation reflects node affinity and taints the Pod cannot tolerate. These fields were beta in v1.26 and graduated to stable in v1.33; on older clusters, verify the feature gate and API schema before applying them. Nodes missing any topology key used by the constraints are bypassed, so they neither contribute to the calculation nor serve as placement targets for those constraints.
Account for autoscaling, rollouts, and storage
The scheduler discovers domains from existing nodes; it does not know every zone that a cloud node group could create. In an autoscaled cluster, a zone with zero nodes may therefore be absent from spread calculations. If hard spread is expected to trigger scale-out into an empty zone, confirm that the node autoscaler understands the Pod constraints and the provider’s zone inventory, or keep suitable capacity represented in each required domain.
During a Deployment rollout, old and new ReplicaSets can coexist. Make sure the selector counts the population you intend: combining revisions may preserve aggregate balance while leaving the new version concentrated, whereas separating revisions can temporarily make each revision harder to place. The beta matchLabelKeys field, enabled by default since Kubernetes 1.27, can include pod-template-hash so the spread calculation is revision-specific. In Kubernetes 1.34 and later, the API server explicitly merges those matching key-value labels into labelSelector; earlier versions handled the field implicitly. Verify the target version and avoid directly editing a Pod label referenced by matchLabelKeys, because that does not update the merged selector.
HPA scale-up creates more Pods subject to the same placement limits; scale-down can leave an uneven distribution because existing Pods are not relocated. A PodDisruptionBudget controls eligible voluntary evictions, not zone placement or involuntary zone loss. For StatefulSets, zonal PV affinity and CSI provisioning can narrow the eligible domains: verify storage topology and binding behavior together with the spread rule before promising cross-zone recovery.
Verify actual placement and failure behavior
Save the first manifest as topology-demo.yaml, then validate and inspect its placement:
kubectl apply --dry-run=server -f topology-demo.yaml
kubectl apply --dry-run=server -f strict-three-zones.yaml
kubectl apply -f topology-demo.yaml
kubectl rollout status deployment/topology-demo --timeout=5m
kubectl get nodes -L topology.kubernetes.io/zone,kubernetes.io/hostname
kubectl get pods -l app=topology-demo -o wide
For a Pending replica, use kubectl describe pod and recent scheduler Events to distinguish a skew rejection from insufficient CPU, taints, required affinity, quota, or volume topology. Check the labels on the candidate nodes and calculate counts only for Pods matched by the constraint selector. A successful server-side dry run validates API acceptance, not that the cluster has enough capacity or failure domains.
Production acceptance should prove the policy, not just the manifest:
- With three labeled, eligible zones and adequate capacity, the six-Pod example reaches Ready and matching zone counts differ by no more than one.
- Host-level spread is evaluated as a preference, not treated as a guarantee; verify observed node placement under realistic resource pressure.
- In a disposable environment, remove capacity from one zone and confirm the strict example’s Pending/placement behavior matches the service’s recovery objective. For
minDomains: 3, verify that reduced-domain Pending replicas are intentional and alertable. - Test scale-up, scale-down, rollout surge, node drain, and a zonal-volume workload where applicable. Recheck Pods, node labels, events, PDB status, and application health after each operation.
Topology spread is a precise scheduling input, not an automatic multi-zone architecture. The failure-domain labels, selector population, eligible capacity, storage model, autoscaler, and disruption policy must all describe the same recovery plan.
Related:
- How the Kubernetes Scheduler Actually Places Workloads
- How to Configure Pod Disruption Budgets in Kubernetes
Sources:
- Kubernetes documentation: Pod Topology Spread Constraints
- Kubernetes API reference: Pod v1
- Kubernetes documentation: Pod disruptions
- Kubernetes documentation: Horizontal Pod Autoscaling
- Kubernetes documentation: StatefulSets
- Kubernetes documentation: Persistent Volumes
- Kubernetes documentation: Storage Classes