Kubernetes DaemonSet Rollouts: maxUnavailable, maxSurge, and Node Coverage
Plan DaemonSet updates across eligible nodes with OnDelete, maxUnavailable, maxSurge, readiness gates, capacity checks, and rollback evidence.
A DaemonSet rollout is a node-by-node change to software that often provides infrastructure for other workloads: logging, monitoring, storage, networking, device management, or security. Unlike a Deployment, a DaemonSet targets eligible nodes rather than a fixed replica count. An update that looks small in a manifest can therefore interrupt a service on a large fraction of the cluster if its availability budget, node selection, readiness signal, or surge capacity is misunderstood.
Plan the rollout around which nodes should run the agent, what the agent provides to each node, and how much overlap the node can support. maxUnavailable trades rollout speed for the number of nodes temporarily without an available DaemonSet Pod. maxSurge trades additional per-node resource use for an opportunity to start and validate the replacement before removing the old Pod. Neither field proves that the application behind the Pod is functioning correctly; readiness and node-level acceptance criteria must reflect the actual service.
Choose the update strategy
DaemonSets support two update strategies:
RollingUpdateis the default. A change to the Pod template triggers the controller to replace Pods progressively across eligible nodes.OnDeleteleaves existing Pods running after a template change and replaces them only after they are deleted. This makes the operator responsible for deciding when replacement happens.
OnDelete can be useful when an external process controls host maintenance or staged node replacement, but it can also leave a mixed-version fleet indefinitely. Track the desired template revision and the number of nodes still running an older Pod. Do not interpret a successful kubectl apply as evidence that all nodes have updated.
RollingUpdate is usually more suitable for a declarative rollout with controller-managed progress. It uses rollingUpdate.maxUnavailable, optional maxSurge, and minReadySeconds to control how replacement proceeds. The API defaults maxUnavailable to 1, maxSurge to 0, and minReadySeconds to 0. Set these explicitly when the default is not an acceptable production policy.
Set a disruption and capacity budget
This excerpt configures a DaemonSet to make a new Pod available before removing the old one, with at most one surging node at a time and no planned unavailable replicas. It requires enough spare resources and compatible host-level behavior to run both Pods briefly on the same node:
spec:
minReadySeconds: 30
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
maxSurge can be an integer or a percentage of the desired DaemonSet Pod count; percentage values round up to at least one. maxUnavailable can also be an integer or percentage and defaults to one. maxUnavailable cannot be zero when maxSurge is zero, so a configuration that asks for neither an unavailable Pod nor a surge Pod is invalid. Verify the API server accepts the intended values on the cluster release in use before committing the rollout policy.
With maxSurge: 0, the controller normally removes an old Pod before creating its replacement, bounded by maxUnavailable. This keeps the per-node Pod count down but introduces a service gap on each updated node. With a positive surge, the old and new Pods may coexist on a node while the new one becomes available. That can double CPU, memory, disk, network, or host resource consumption for a node during the update. If the new Pod becomes unavailable, the controller may create a replacement without being constrained by the normal surge limit to make progress; operators should plan for that exceptional overlap too.
Choose values from the service’s failure model, not a desired rollout duration alone. A log shipper might tolerate a brief gap if its local buffer is durable. A node networking agent may be part of the node’s ability to become Ready, so removing the working Pod too early can create a bootstrap loop. A CSI node plugin or device agent may be constrained by exclusive host paths or sockets that prevent a second instance from starting safely.
Make availability reflect actual readiness
minReadySeconds is the minimum time a new Pod must remain Ready without a container crash before it counts as available. The default is zero, which lets an immediately Ready Pod count as available. Increase it when the agent has a known warm-up interval, but do not use it as a substitute for a meaningful readiness probe. A process-level probe that succeeds while the agent has not registered with the host subsystem can still let the rollout advance prematurely.
Readiness should answer whether the node-local function is usable. For a log collector, this could mean the agent has loaded its configuration and can reach its local ingestion path; for a CNI component, the true criterion is driver-specific and can involve node network readiness. Avoid probes that require an external dependency unrelated to whether the local DaemonSet Pod can provide service, unless that dependency is part of the actual readiness contract.
Observe both Pod status and DaemonSet status. Useful fields include desiredNumberScheduled, currentNumberScheduled, updatedNumberScheduled, numberReady, numberAvailable, and numberUnavailable. These counts are node-oriented and may change when node labels, taints, readiness, or pool membership changes during a rollout. A rollout can be blocked because a Pod cannot schedule, an image cannot be pulled, a resource request no longer fits, or a readiness condition never becomes true.
Define and inspect the eligible node set
A DaemonSet creates one Pod for each eligible node according to its selector, affinity, and tolerations. Before an update, inspect the effective node inventory rather than assuming the DaemonSet should cover every worker. A new node label or taint can change the desired set during the rollout. A Pod that is intentionally limited to GPU workers, control-plane hosts, or a canary pool should be measured against that subset.
kubectl get daemonset node-agent -n kube-system -o wide
kubectl get daemonset node-agent -n kube-system -o jsonpath='{.status.desiredNumberScheduled}{" desired, "}{.status.updatedNumberScheduled}{" updated, "}{.status.numberAvailable}{" available, "}{.status.numberUnavailable}{" unavailable\n"}'
kubectl get pods -n kube-system -l app=node-agent -o wide
kubectl get nodes -L node-pool,topology.kubernetes.io/zone
DaemonSet Pods receive special tolerations for several node conditions and can be scheduled on unschedulable nodes. This is intentional for node-level infrastructure, but it can surprise operators during maintenance. Confirm the effective Pod tolerations and node selectors; do not infer that cordoning a node has prevented the DaemonSet from running there.
If an agent must be rolled out to only a subset of nodes first, use an explicit canary pool or a carefully reviewed node label selector. Avoid running two full DaemonSets with overlapping selectors against the same node when both write the same host paths, bind the same host ports, or manage one host-level service. A staged rollout must preserve unambiguous ownership of each node-local function.
Account for DaemonSet-specific surge risks
Surge is attractive because the old agent can remain until the new one is Ready, but readiness is not the only constraint. Two Pods may conflict over a host port, a Unix socket, a writable hostPath, a device handle, a kernel module, or a singleton process. Even if both Pods can start, duplicated log collection or metric scraping can create duplicate data. Validate coexistence in a representative node before setting a positive surge.
Resource planning should consider the worst node, not just the cluster average. A percentage surge rounds up and the DaemonSet API allows a new Pod for a node whose old Pod has become unavailable even when the normal surge limit would otherwise be exceeded. Busy nodes can evict unrelated workloads if the extra Pod cannot fit. Review requests, limits, priority, eviction behavior, and disruption to colocated critical workloads before increasing surge.
maxUnavailable has a different cost. An update with a large unavailable allowance can remove the node service from many eligible nodes at once. The percentage is calculated against the desired number of DaemonSet Pods at the start of the update, with percentage conversion rounded up. Model the absolute number at the current fleet size. A seemingly small percentage can still affect multiple nodes in a large cluster, while an absolute value of one may be deliberately conservative but slow.
Do not confuse DaemonSet update limits with a Deployment’s maxSurge behavior or a PodDisruptionBudget. A DaemonSet update is governed by its own strategy fields. Plan voluntary node maintenance and controller-driven updates separately, and verify which mechanism performs the Pod termination in the incident or change path.
Execute, watch, and pause with evidence
Before applying a template change, capture the current DaemonSet revision, healthy Pod placement, node inventory, and a known-good image or configuration revision. Confirm that the rollback artifact remains available and that the change is reversible. Apply through the repository or release controller that owns the DaemonSet rather than making a one-off live edit that GitOps will immediately overwrite.
kubectl rollout status daemonset/node-agent -n kube-system --timeout=10m
kubectl rollout history daemonset/node-agent -n kube-system
kubectl describe daemonset/node-agent -n kube-system
kubectl get events -n kube-system --sort-by=.lastTimestamp
Watch Pod readiness per node, not just the aggregate updatedNumberScheduled. A completed rollout is not proof that the node service behaves correctly. Monitor the agent’s own health, host-level integration, data loss or duplication, CPU/memory overhead, and user-facing service signals. For a node networking or storage component, run a canary workload that exercises the affected capability before proceeding to the next fleet segment.
Pause or revert when the new Pod fails readiness, old Pods disappear too early, duplicate processing appears, host resources saturate, or the eligible node set changes unexpectedly. For a rollback, use the established source-of-truth workflow and verify both the PodTemplate revision and the effective Pods on nodes. A template rollback may itself trigger a new DaemonSet rollout; it is not instantaneous restoration.
For OnDelete, keep an explicit per-node replacement plan. Delete Pods in bounded batches only after validating that the new template is correct and the node can tolerate the transition. Track every node that still runs the old revision and define a deadline for completing the manual process so that a partial update does not become permanent.
Production acceptance checklist
- The update strategy and
maxUnavailable/maxSurgematch the service’s node-level availability objective. - The actual eligible node count and selectors are known, including canary and control-plane pools.
- Readiness verifies the node-local function, and
minReadySecondscovers meaningful warm-up rather than masking a bad probe. - Surge coexistence has been tested for ports, host paths, device access, duplicate work, and worst-node capacity.
- Image availability and rollback revision are confirmed on cold/replacement nodes.
- The rollout is observed per node and through the agent’s functional signals, not only controller counters.
- A pause, rollback, or
OnDeletemanual sequence has been rehearsed and has a clear owner.
A production-grade DaemonSet rollout is a controlled migration of node-local capability. The right strategy minimizes both the period without a functioning agent and the risk of running two incompatible instances on one node. Make the budget explicit, watch each node transition, and stop on functional evidence rather than waiting for the controller to report a green summary.
Related:
- Kubernetes Deployment Rollouts: Progress, Availability, and Revision History
- Kubernetes Graceful Node Shutdown: Kubelet Budgets, Priority, and Limits
Sources: