Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

Kubernetes Pod Readiness Gates: Extending Traffic Eligibility with Custom Conditions

Use custom Pod conditions to gate Kubernetes readiness on external prerequisites, with controller design, Service routing, rollout safety, and verification.

A readiness probe answers whether a container reports itself ready. It cannot, by itself, prove that an external control plane has programmed a route, attached a storage path, completed a security handshake, or registered the instance with a system outside Kubernetes. A Pod readiness gate adds a second, explicit input to the Pod’s overall Ready condition: a custom Pod condition that another controller owns.

Pod readiness gates have been stable since Kubernetes 1.14. They add an external condition alongside the normal requirements that the node and containers are ready; they do not replace those checks.

This is a coordination mechanism, not another probe type. The kubelet evaluates the condition named in the Pod specification, but an application operator or controller must write that condition. A missing condition is treated as false. The Pod can have healthy, ready containers and still remain not Ready until every configured readiness gate is True.

Model the gate as an external readiness contract

Use a gate only when a meaningful readiness fact is owned outside the container or cannot be established by the Pod’s own readiness endpoint. Examples include an infrastructure controller confirming that a backend has been registered with a load balancer, a service-mesh controller confirming that an identity or route is installed, or an operator confirming that a node-local resource is attached and usable. These examples are design patterns, not built-in Kubernetes integrations: the external system and the controller that observes it must be deployed and operated separately.

The contract has three parts:

  1. The Pod template declares a stable condition type in spec.readinessGates.
  2. A responsible controller observes the Pod and the external prerequisite, then updates the matching Pod condition.
  3. Kubernetes combines that condition with container readiness when computing Pod Ready. A condition that is absent, False, or Unknown does not satisfy a gate.

The condition type must use a valid Kubernetes label-key format. Prefer a domain controlled by the organization, such as delivery.example.net/route-programmed, rather than an unqualified name that can collide with another integration. Treat the condition type as an API: document its owner, the exact evidence needed for True, and what should make it return to False.

Declare the custom condition in the workload template

The following Deployment fragment is illustrative. Replace the registry, image, port, and probe path with values from the application. The readinessGates field belongs in the Pod template, not on a Service or Deployment status object.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: checkout-api
  namespace: production
spec:
  replicas: 3
  selector:
    matchLabels:
      app: checkout-api
  template:
    metadata:
      labels:
        app: checkout-api
    spec:
      readinessGates:
        - conditionType: delivery.example.net/route-programmed
      containers:
        - name: api
          image: registry.example.net/checkout-api:reviewed
          ports:
            - name: http
              containerPort: 8080
          readinessProbe:
            httpGet:
              path: /ready
              port: http
            periodSeconds: 5
            timeoutSeconds: 2
            failureThreshold: 2

The readiness probe and custom gate express different facts. The probe lets the kubelet ask the process whether it can serve its own work. The gate lets a controller report that an external prerequisite has been satisfied. A True gate does not make a broken process healthy, and a successful probe does not prove that the external system can reach the Pod.

In this example, an application controller should set delivery.example.net/route-programmed to True only after it has observed evidence that the intended route or backend registration is effective for this specific Pod. Merely submitting a configuration request to another API is not sufficient if that API processes changes asynchronously. If the system can report Pending, Failed, or Unknown, map those states deliberately: only confirmed success should produce True.

Implement the condition writer as a controller

The readiness gate is passive. It does not call the external API, retry a route operation, or decide when a stale registration should be removed. Those responsibilities belong to the controller. Use a reconciliation loop that can safely run more than once: read the current Pod by UID, inspect the authoritative external state, and patch the matching condition only when the observed state differs from the desired condition.

Patch the Pod status subresource rather than changing spec or replacing the full status object. Kubernetes reserves built-in conditions and container status fields for system components. A controller should update only its own condition type, preserve conditions owned by other components, and set a useful reason, message, and transition time. Give the controller narrowly scoped RBAC for reading the relevant resources and patching Pod status in the required namespaces.

Bind external registrations to the Pod UID, not only to namespace and name. Pod names can be reused after deletion; a delayed callback for an old Pod must not mark a replacement Pod Ready. On deletion or replacement, reconcile cleanup of the old external registration and stop writing status for the old UID. If the external system is unavailable, do not leave an old True condition indefinitely: define whether the evidence expires, whether the controller rechecks it, and how uncertainty is represented.

Avoid a circular dependency. If the external registration controller watches only ready Pods, but the readiness gate waits for that registration, neither side can make progress. The controller must observe Pods before they are Ready, while ensuring its own side effect is safe and idempotent. Likewise, do not make a route controller set the condition based on a cached success response without verifying that the configuration still belongs to the current Pod.

Understand what Services and rollouts observe

For an ordinary Service, Pod readiness affects whether its endpoint is considered eligible for normal traffic. EndpointSlices expose readiness-related endpoint conditions, and consumers should account for serving and terminating state rather than assuming a Pod’s container state is the complete routing picture. With spec.publishNotReadyAddresses enabled, EndpointSlice readiness has different behavior; do not expect a custom gate to keep such endpoints out of every consumer’s view.

The gate also affects workload availability. A Deployment counts a Pod as available only after it is Ready for the configured minimum-ready interval. If the external condition never becomes True, a rollout may leave new Pods unready, retain old replicas, or eventually report a progress failure. That is often the intended fail-closed outcome, but it can reduce capacity. Size rollout surge and spare capacity for the time needed by the external integration, and alert on long-lived false or missing gate conditions.

Readiness gates are not a substitute for safe shutdown. During termination, EndpointSlice conditions and endpoint consumers have their own handling for terminating Pods. The controller that owns an external registration should remove it when the Pod is draining or deleted, but should not assume that flipping a custom condition instantly drains every connection. Validate the behavior of the actual ingress, proxy, mesh, and external load balancer.

Common failure modes

  • A condition type typo in spec.readinessGates never becomes True because the controller writes a different type. Compare the exact string in the Pod spec and Pod status.
  • The controller updates a Pod by name after a long external operation and accidentally writes a stale success for a replacement Pod. Re-read and compare UID immediately before the status patch.
  • The readiness controller loses permission to patch the status subresource. The Pods remain gated while its logs show authorization errors; verify RBAC for pods/status, not only pods.
  • The external API accepts a request but the data plane has not converged. Report True only at the point specified by the readiness contract, and monitor the difference between accepted and effective state.
  • A gate is added to a Deployment with no plan for old and new replicas. Changing the Pod template creates replacement Pods with the gate; existing Pods do not acquire the new readiness behavior until replaced.
  • A gate depends on a service that itself routes only to Ready Pods from the same workload. This readiness cycle can leave every replica unavailable. Break the cycle with an independent bootstrap path or a different health contract.
  • An operator enables publishNotReadyAddresses for peer discovery and assumes the gate hides the Pod from all clients. Verify each consumer’s behavior because that Service setting intentionally changes readiness publication.

Verify the control path, not just the Pod phase

Inspect the exact custom and Ready conditions, then verify the EndpointSlice produced for the Service:

kubectl get pod POD_NAME -n production -o yaml
kubectl describe pod POD_NAME -n production
kubectl get endpointslice -n production \
  -l kubernetes.io/service-name=checkout-api -o yaml
kubectl rollout status deployment/checkout-api -n production
kubectl get events -n production --sort-by=.metadata.creationTimestamp

The expected sequence is: the Pod is created with the gate; the application becomes container-ready; the custom controller observes the external prerequisite; the controller patches the Pod condition; the kubelet reports Ready; and the service endpoint becomes eligible according to the Service and EndpointSlice rules. Test the negative path too: delay or fail the external prerequisite and confirm that the Pod is not published as ready, the rollout alerts, and existing healthy replicas remain available.

Before production, define an objective timeout for the external prerequisite, the alert for a gate that remains false, the recovery or rollback action, and how the controller cleans up registrations. Gate correctness is a safety property only when its status reflects current evidence. A stale True can send traffic to an unusable backend; a stale False can block a healthy rollout. Monitor both conditions and reconciliation age.

Related:

Sources:

Comments