Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

Kubernetes Owner References and Garbage Collection: Deletion Is a Workflow

Understand owner references, cascading deletion, orphaning, and finalizers so Kubernetes cleanup does not remove or strand dependent resources unexpectedly.

Kubernetes object deletion is not always an immediate removal of one API record. The garbage collector can follow owner references and delete dependent objects, a controller may hold an object with a finalizer while cleanup runs, and an API client may choose to orphan dependents deliberately. Understanding those relationships is essential before deleting a Deployment, custom resource, namespace, or storage object.

Owner references express lifecycle ownership: a dependent object points to the API identity of its owner. Labels and selectors are different. A controller may use labels to discover candidates, but an owner reference says which object’s deletion should affect the dependent. For example, a Service’s EndpointSlices can have both a Service-name label for lookup and an owner reference for cleanup. Deleting by label and deleting by owner are not equivalent operations.

Inspect ownership as an API relationship

An owner reference includes the owner’s API version, kind, name, UID, and optional flags such as controller and blockOwnerDeletion. The UID matters: deleting an object and recreating another object with the same name produces a different API identity. A dependent referencing the old UID is not automatically owned by the new object just because their names match.

kubectl get replicaset catalog-7c8bd8c7f9 -n production -o yaml
kubectl get deployment catalog -n production -o yaml
kubectl get pods -n production -l app=catalog -o yaml

Inspect metadata.ownerReferences, metadata.uid, metadata.deletionTimestamp, and metadata.finalizers. A Pod may be controlled by a ReplicaSet, while that ReplicaSet is controlled by a Deployment. Removing an intermediate owner can therefore change which controller is responsible for a child object. Do not patch owner references by hand as a routine repair: the owning controller may reconcile them back, and a wrong UID or scope can result in orphaning or blocked cleanup.

Namespaced dependents can refer to an owner in the same namespace or to a cluster-scoped owner. Cross-namespace owner references are invalid by design. Cluster-scoped dependents cannot validly reference a namespaced owner. In recent Kubernetes versions, the garbage collector reports a warning Event with reason OwnerRefInvalidNamespace for these invalid relationships. Investigate those events rather than repeatedly deleting an object that cannot be collected through the expected owner link.

Choose the deletion propagation mode deliberately

Foreground cascading deletion keeps the owner visible while the garbage collector removes dependents that block deletion. The owner receives a deletion timestamp and the foregroundDeletion finalizer; only after the required dependents are gone is the owner removed from the API. This is useful when the caller needs to wait for dependent cleanup before treating the owner as deleted.

Background cascading deletion removes the owner object promptly and cleans up its dependents asynchronously. Background is the default propagation policy. A query that no longer finds the owner does not prove that all children have already disappeared. If downstream cleanup must complete before another action, wait for the dependents or use a workflow that observes their state rather than assuming the API request was synchronous.

Orphan propagation deletes the owner but leaves the dependent objects. This is intentional only when another controller or operator is prepared to adopt or clean up those objects. Orphans can continue consuming CPU, storage, IP addresses, and cloud resources. Record the intended next owner and test adoption behavior before using this mode on production resources.

# Inspect the dependency graph before choosing a propagation policy.
kubectl get deployment catalog -n production -o yaml
kubectl get replicasets,pods -n production -l app=catalog -o wide

# If deletion is approved, choose one documented mode explicitly.
kubectl delete deployment catalog -n production --cascade=foreground

The command is an operational example, not a recommendation to delete a live Deployment without review. Foreground deletion can take longer or remain pending when a dependent finalizer is stuck or the garbage collector cannot list a resource. Background mode can make parent disappearance appear successful before child cleanup finishes. Observe events and object state with a deadline and escalation path.

Finalizers provide a cleanup checkpoint

A finalizer is a key in metadata.finalizers that keeps an object in a terminating state until a controller performs its cleanup and removes the key. When deletion begins, the API server sets metadata.deletionTimestamp; the responsible controller observes that state, releases external resources, and then removes its finalizer. Garbage collection and finalizers work together, but they solve different parts of the lifecycle.

A finalizer that never clears can leave an object terminating indefinitely. Diagnose which controller owns the finalizer, whether that controller is running, and whether its cleanup call is failing. Check controller logs, events, required API permissions, network connectivity, and the external resource state. If the original controller is gone, identify and clean the external resource through its owning system before deciding whether removing the finalizer is safe.

Do not remove finalizers simply to make kubectl get look clean. A finalizer may be the only record that cloud infrastructure, storage, DNS, or another external dependency still needs cleanup. Force-removing it can orphan billable or stateful resources. Conversely, waiting forever without understanding the finalizer can block namespace deletion and operational recovery. Capture the object YAML and controller evidence first, then follow the operator’s documented recovery procedure.

The blockOwnerDeletion field affects foreground deletion behavior when the owner has a foregroundDeletion finalizer. Only dependents carrying the relevant flag and visible to the garbage collector cache can block owner deletion. This is not a generic “delete children first” flag. Understand which controller sets it and why before changing it.

Keep garbage collection separate from other cleanup mechanisms

Kubernetes uses “garbage collection” for multiple mechanisms. API object garbage collection follows ownership and controller-specific rules. The kubelet also removes unused containers and images according to node-level policies. PersistentVolume reclaim policy controls what happens to backing storage after claim release. TTL controllers can clean up completed Jobs. These processes have different owners and safety properties; seeing a resource disappear does not identify which mechanism removed it.

For example, deleting a PVC does not necessarily mean its storage backend is deleted immediately; the StorageClass reclaim policy and CSI driver affect that lifecycle. Likewise, garbage collection of a Deployment’s Pods does not clean arbitrary cloud resources created by a custom operator unless that operator modeled the ownership and finalizer behavior. Before changing cleanup, identify the API object, controller, external resource, and retention requirement involved.

Kubernetes recommends that each API client’s cleanup be designed around explicit ownership and lifecycle. Controllers should use idempotent reconciliation: repeated cleanup attempts must be safe, and a temporary API or cloud-provider failure should not permanently lose the information needed to retry. Operators should emit status and events that identify which prerequisite or external operation prevents cleanup.

Run a deletion drill before relying on the policy

In a disposable namespace, create a representative resource graph, observe owner references, and exercise foreground, background, and orphan behavior. Include one dependent with a simulated cleanup delay and one finalizer controlled by a test controller. Record timestamps for the delete request, owner disappearance, child disappearance, and external cleanup completion. This exposes whether the automation’s wait condition matches the actual lifecycle.

For production procedures, preserve the owning objects’ YAML and UIDs, understand selector and ownership links, and decide how to handle terminating objects before issuing a delete. If a resource is unexpectedly stuck, gather events and controller logs before editing finalizers. Validate the resulting external state after cleanup, because a clean API list is not proof that external resources were deleted.

The operational rule is simple: deletion is a distributed workflow. Owner references define relationships, propagation policy defines how Kubernetes starts the cascade, finalizers provide controller-managed checkpoints, and external systems may have their own asynchronous cleanup. A safe runbook makes each stage observable and provides a recoverable response when any stage fails.

When cleanup is stuck, inventory dependents by owner UID rather than just by matching labels. Labels can include unrelated objects or omit an object managed by a different controller. Compare owner and dependent namespaces, API group and kind, UID, deletion timestamps, finalizer strings, and recent Events. If the owner is absent but a child remains, determine whether it was intentionally orphaned, has another owner, or has an invalid reference before changing it.

For operators, expose cleanup progress in the custom resource’s status and make retries safe. A controller should not remove its finalizer until its required external cleanup has succeeded or a documented terminal outcome has been recorded. Add a bounded retry/backoff strategy and an alert for objects that remain terminating beyond the expected service-level window. This turns finalizer diagnosis from a manual guess into a measurable lifecycle condition.

Related:

Sources:

Comments