Kubernetes Controller Reconciliation: Idempotent Desired-State Loops
Design resilient Kubernetes controllers around level-based reconciliation, stale caches, retryable side effects, finalizers, status, and operational signals.
A Kubernetes controller is a control loop: it compares declared intent with observed state and makes changes that move the system toward that intent. A reconcile invocation is not a transaction or an instruction to “execute this event once.” Design for retries, restarts, and observations that lag behind writes.
Reconcile levels, not event payloads
A controller-runtime request identifies an object by namespace and name; it contains neither the triggering event nor a snapshot. Watches enqueue keys, and the reconciler reads the latest primary object and dependencies. Multiple updates may collapse into one request, so never depend on observing every intermediate change.
The usual entry point is intentionally small:
var widget examplev1.Widget
if err := r.Get(ctx, req.NamespacedName, &widget); err != nil {
return ctrl.Result{}, client.IgnoreNotFound(err)
}
Derive desired children from the current spec and compare them with current children and, when applicable, the external provider. A delete notification is handled by observing that an owned object is absent; an update notification by reading its latest state. Startup and retries can repair drift without replaying an event log.
Dependent-object events do not automatically become requests for every controller that might care about them. Configure watches and map child keys back to the primary object, commonly through an owner reference or an explicit mapping function. The work queue de-duplicates identical keys, so several changes can arrive as one reconcile. When a relationship changes from one parent to another, enqueueing only the new owner can leave the old owner’s external or child state behind; reconcile both sides or design cleanup to be discoverable from durable state. Event predicates can reduce noise, but a predicate that filters a meaningful state transition also removes that repair trigger.
Make every mutation safe to repeat
Assume the same key can be reconciled again after success, after an API conflict, or after a timeout whose outcome is unknown. Use deterministic child names, owner references, and stable external identities so a retry can find what the prior attempt created. Update only fields the controller owns, and write only when the desired and observed values differ. For shared objects, establish field ownership deliberately rather than continually replacing another actor’s fields.
External APIs make the failure boundary especially important. A provider may commit while its response is lost, or the controller may crash before updating Kubernetes status. The next pass must find or safely upsert the same resource. Use a provider idempotency key or deterministic identifier; include a stable revision when an operation is specific to one spec generation. The reconcile loop itself supplies no atomic transaction across Kubernetes and the provider, so it cannot promise exactly-once side effects.
Account for cache lag and queue retries
The default controller-runtime client reads from a local cache and writes to the API server; a read immediately after a write may still be stale. Do not interpret that view as proof a write failed and create a randomly named duplicate. Deterministic names and optimistic concurrency make repeats detectable; otherwise, persist enough operation identity to discover or retry safely.
Return an error for a transient failure and let the controller’s rate-limited retry policy provide backoff. Use a timed requeue when the next useful action depends on elapsed time, such as polling an asynchronous provider operation:
return ctrl.Result{RequeueAfter: 30 * time.Second}, nil
Choose an interval that respects provider limits and the user-visible freshness target. Avoid immediate requeue loops for permanent validation errors; report those as a condition and wait for a meaningful spec change or a deliberate periodic check. A returned retry or timer is a scheduling hint, not a durable work record, so important external operations still need recoverable identity and observable progress.
Treat finalizers and status as part of the protocol
If a custom resource owns an external resource that Kubernetes cannot garbage-collect, add a qualified finalizer before provisioning it. When deletion is requested, the API sets deletionTimestamp and retains the object while finalizers remain. The controller should idempotently delete or confirm absence of the external resource, then remove only its own finalizer. Since Kubernetes does not allow adding a new finalizer once deletion has started, adding it after the external side effect creates a leak window.
Publish observed outcomes in status, not by rewriting spec. Conditions should explain what is ready, progressing, or blocked and give a useful reason and message. If the API defines observedGeneration, use it to distinguish status for the current spec from stale status for an earlier generation. Avoid status writes when values have not changed: unnecessary updates can trigger more reconciliations and create noisy API traffic.
Status is a report of the controller’s latest observation, not a lock or a promise that the underlying system cannot change. Include the generation the controller evaluated when the API supports it, and avoid setting a Ready condition for a different spec generation. Kubernetes update conflicts are normal when another actor changes the object between read and write; fetch fresh state and recompute the owned fields instead of replaying a stale whole-object update. Keep external calls outside a generic API conflict retry unless they are idempotent, because repeating a callback can repeat the real-world side effect even when the Kubernetes write did not happen.
Measure convergence, not just process health
Emit structured logs with resource identity, generation, reconcile ID, action, and provider request ID, while keeping credentials and sensitive payloads out of logs. Track reconcile duration and errors, retry frequency, provider latency and throttling, and how long resources remain out of sync. Alert on actionable conditions such as a finalizer stuck beyond its service objective or a resource that remains unready; a live controller process alone does not prove the loop is making progress. Avoid resource names as metric labels because per-object cardinality grows with the workload.
Test duplicate triggers, coalesced updates, API conflicts, stale reads, provider timeouts after successful commits, restarts between side effect and status update, missing children, and deletion during provisioning. Verify repeated passes converge, cleanup is safe to repeat, and status exposes blocked dependencies. Reconciliation promises eventual convergence under retries and partial failure, not one execution per event.
Related:
- Kubernetes CronJob Execution Semantics: Designing Reliable Scheduled Work
- Kubernetes Server-Side Apply: Field Ownership, Conflicts, and managedFields
Sources: