Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

Kubernetes Deployment Rollouts: Progress, Availability, and Revision History

Read Deployment rollout conditions correctly, size rolling-update limits, and understand what progress deadlines and rollback actually change.

A Kubernetes Deployment rollout is a controller-managed transition between Pod templates. Updating the template creates a new ReplicaSet; the Deployment controller then scales old and new ReplicaSets according to the strategy, readiness, and availability constraints. kubectl apply accepting a manifest is only the start of the process. A production release needs a bounded rollout, observable progress, an application-level success check, and a rollback decision that accounts for what Kubernetes actually stores in a revision.

The default rolling-update strategy gradually replaces old Pods. The maxSurge setting bounds how many extra Pods can be created above the desired replica count, while maxUnavailable bounds how many desired Pods may be unavailable during the update. Values can be integers or percentages. They trade temporary capacity for rollout speed and availability; they should be sized against cluster capacity, startup latency, and the application’s concurrency behavior.

Use a rollout budget that matches the service

For a small stateless API, a rolling update might create one extra replica while keeping all desired replicas available. For a large memory-heavy workload, that extra replica may not fit on the cluster. For a service with a warm-up requirement, the controller must wait for readiness and optionally for minReadySeconds before treating a new Pod as available. Select limits based on measured rollout behavior rather than copying a generic YAML fragment.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: catalog
  namespace: production
spec:
  replicas: 4
  minReadySeconds: 15
  progressDeadlineSeconds: 900
  revisionHistoryLimit: 8
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  selector:
    matchLabels:
      app: catalog
  template:
    metadata:
      labels:
        app: catalog
    spec:
      containers:
        - name: catalog
          image: example.invalid/catalog:2026.10.03
          ports:
            - name: http
              containerPort: 8080

The image name is intentionally illustrative. Replace it with an immutable tag or digest appropriate to the release process, and ensure the cluster can pull it. maxUnavailable: 0 preserves the desired available count during a healthy rollout, but it can stall if the cluster lacks spare capacity or new Pods cannot pass readiness. A budget that promises zero unavailability does not create capacity; it creates a clear constraint that the scheduler and controller must satisfy.

minReadySeconds requires a new Pod to remain Ready for the configured duration before the Deployment considers it available. This can reduce promotion of a Pod that fails immediately after first readiness, but it does not replace meaningful readiness checks. progressDeadlineSeconds controls how long the controller waits without progress before reporting a ProgressDeadlineExceeded condition. The controller continues retrying; Kubernetes does not automatically roll back the Deployment when the deadline is exceeded.

Observe desired state and actual progress

Use rollout status with an explicit timeout, then inspect Deployment conditions, ReplicaSets, Pods, events, and application health. A Progressing=True condition can mean the rollout has made progress, but it is not by itself proof that every replica is updated and available. Confirm the updated replica count, ready replicas, available replicas, and the Pod template image.

kubectl rollout status deployment/catalog -n production --timeout=15m
kubectl get deployment catalog -n production -o wide
kubectl get replicasets -n production -l app=catalog
kubectl get pods -n production -l app=catalog -o wide
kubectl describe deployment catalog -n production

If rollout status times out, capture the Deployment’s status and events before making another change. New Pods may be pending due to resource requests, pull errors, image architecture mismatch, admission rejection, or scheduling constraints. They may start and then fail readiness, or become Ready but violate minReadySeconds stability. Each case needs a different fix; increasing the deadline can hide a slow failure without correcting the cause.

Pause and resume are useful for staged workflows, but a paused Deployment is not a health gate. A pipeline can pause after a low-risk change, run synthetic checks, then resume. Make sure the pipeline detects a failed check and does not leave the controller paused indefinitely. kubectl rollout status should be run after resume as well as before a promotion decision.

Understand revision history and rollback limits

Deployment revisions are created when the Pod template changes. Scaling the Deployment does not create a new revision, which allows manual or autoscaled replica changes without polluting rollout history. Rolling back restores the previous Pod template; it does not necessarily restore other Deployment settings such as the current replica count. Record release metadata in a way that allows operators to identify the code and configuration represented by each revision.

kubectl rollout history deployment/catalog -n production
kubectl rollout history deployment/catalog -n production --revision=12

# Use only after confirming the target revision and impact.
kubectl rollout undo deployment/catalog -n production --to-revision=12
kubectl rollout status deployment/catalog -n production --timeout=15m

The revisionHistoryLimit controls how many old ReplicaSets are retained. A value that is too small can remove rollback targets sooner than the incident response horizon; a value that is too large retains more API objects. Deleting ReplicaSets manually to “clean up” can remove useful history. If release artifacts or configuration references have their own retention policies, align them with Deployment rollback history.

Rollback is not a database migration strategy. A previous application version may not understand data written by the new version, and a template rollback cannot reverse an external side effect or restore a deleted ConfigMap. Use expand-and-contract schema migrations, backward-compatible configuration, and an explicit rollback test where stateful dependencies are involved.

A rollout is not a traffic canary by itself

RollingUpdate controls how Pods are replaced, not how requests are distributed between versions at an application or user cohort level. During a rollout, old and new Pods may both receive Service traffic. maxSurge and maxUnavailable shape replica counts, but they do not provide a percentage-based canary, metric analysis, or automatic rollback. Use a dedicated traffic-management strategy if a release requires those semantics.

Readiness determines whether a Pod should receive normal Service traffic. A new version that becomes Ready too early can receive production requests before cache warm-up or dependency checks finish. A readiness check that depends on a noncritical downstream service can instead keep every new Pod unavailable during a transient dependency incident. Define readiness around the ability to serve the intended request class and test it under partial failure.

PodDisruptionBudgets are primarily for voluntary disruptions such as node drains; they do not replace the Deployment’s rollout strategy fields. Set maxSurge and maxUnavailable for Deployment rollout behavior, and use disruption budgets for the supported disruption scenarios they govern. Verify the exact controller behavior for the Kubernetes version and managed platform in use.

Test rollout failure and recovery

In a staging cluster, test a successful release, an image pull failure, a Pod that remains unready, and a rollout that exceeds the progress deadline. Confirm the pipeline reports the failure, captures diagnostics, and stops promotion. Then test rollback to the previous template and verify that the old image and required configuration still exist. Include capacity saturation if maxSurge requires spare resources.

Record rollout identifiers, image digests, configuration version, start time, and health results. Alert on a false Progressing condition, low availability, and prolonged rollout duration. Do not automatically issue rollout undo solely because the progress deadline elapsed: a temporary scheduling constraint may resolve without a rollback, and a rollback could be unsafe after a data migration. Make the recovery policy a deliberate operator decision supported by observable evidence.

A reliable Deployment pipeline treats rollout status as one signal in a larger release contract. The controller can report that the desired template has converged; application metrics, synthetic requests, and dependency checks determine whether the release is actually healthy. Preserve both perspectives so a green Kubernetes condition does not conceal a user-facing regression.

For a release gate, capture a baseline before updating the template: desired and available replicas, error rate, latency distribution, saturation, and the currently serving image digest. After the rollout, compare the same signals over a fixed observation window. A Pod that passes readiness once can still produce elevated error rates under real traffic. Conversely, an external dependency outage can increase errors for both old and new versions, so rollback should be based on evidence that the new revision is responsible rather than on a single uncorrelated alert.

Automated release tooling should retain the exact ReplicaSet revision and artifact digest it promoted. Tags can be moved or deleted, and a revision number alone is meaningful only within one Deployment object’s history. Record namespace, cluster, Deployment UID, revision, template hash, image digest, and configuration version together. This makes a later rollback review reproducible even if a Deployment was deleted and recreated under the same name.

Related:

Sources:

Comments