Skip to content
SRE & DevOpsDeep Dive Published Updated 10 min readViews unavailable

Kubernetes Image Garbage Collection: Kubelet Thresholds, Runtime Storage, and Cache Planning

Configure kubelet image reclamation thresholds and age limits without confusing cached images with Pod eviction, ephemeral storage, or runtime disk layout.

Kubernetes image garbage collection is a node-local kubelet responsibility: it removes images that are not in use when the kubelet’s image manager decides they should be reclaimed. It is neither a registry-retention policy nor a guarantee that an image will remain cached for the next rollout. The distinction matters during node replacement, rapid deployment rollouts, large-scale pulls, and disk-pressure incidents: a healthy Pod can keep its image usable while the same node removes other unused layers, then a later Pod may need to download them again.

Treat the image cache as expendable capacity. Keep authoritative image references in a registry, pin production workloads to immutable digests when reproducibility requires it, and model node storage so that image cleanup is normal housekeeping rather than an emergency command. Do not run an independent runtime cleanup tool against a Kubernetes node: the Kubernetes documentation warns that external garbage collection can remove containers that the kubelet expects to exist.

Understand the two threshold model

The kubelet has high and low image disk usage thresholds. When measured usage rises above the high threshold, image garbage collection runs and removes eligible images until usage falls to the low threshold. The current kubelet configuration API documents defaults of 85 percent for imageGCHighThresholdPercent and 80 percent for imageGCLowThresholdPercent. The high value must be greater than the low value; both are integer percentages from 0 through 100.

These are trigger and target values, not a reservation for each Pod. A five-point gap is hysteresis: it avoids running cleanup continuously at one exact boundary. It does not ensure a fixed number of gigabytes will remain free, because image sizes, filesystem capacity, writable layers, logs, and other consumers vary. A 100 GiB image filesystem with an 80 percent target has different absolute headroom from a 500 GiB filesystem with the same percentage.

The kubelet orders unused images by when they were last used, oldest first, and continues reclamation toward the low threshold. Images currently used by containers are not candidates for this unused-image cleanup. That protects running workloads from this particular mechanism, but it does not make a node’s image cache durable or prevent a later pull from being slow or unavailable.

Configure through the kubelet configuration file

The kubelet configuration API is the maintainable place to set these values. The standalone command-line flags for the image high and low thresholds are deprecated in favor of the configuration file. A platform may manage the file through kubeadm, an image build, a managed node group, or another provisioning system; make the change through that owner instead of editing one node by hand and allowing automation to revert it.

This example changes the high/low thresholds and adds an unused-image age limit. It is an example policy, not a universal recommendation:

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
imageGCHighThresholdPercent: 85
imageGCLowThresholdPercent: 75
imageMinimumGCAge: 2m
imageMaximumGCAge: 168h

Before applying a configuration, inspect the complete node configuration and the distribution’s supported configuration workflow. Kubelet configuration files are not necessarily merged field by field with defaults: for example, Kubernetes documents that changing one evictionHard parameter while omitting others can cause the omitted thresholds to become zero unless defaults are explicitly merged. That warning concerns eviction settings, not image-GC thresholds, but it illustrates why a partial file must not be assumed to preserve every default.

Roll out a configuration change to a representative node pool, confirm kubelet starts with the intended configuration, and observe image pulls, free space, and workload startup latency during the next normal rollout. The exact restart or node replacement procedure is platform-specific. Do not edit a production kubelet file or restart a node solely because the sample above looks reasonable.

Separate minimum and maximum unused age

imageMinimumGCAge is a floor: an image must be unused for at least this long before it can be garbage collected. The documented default is two minutes. It can prevent a just-replaced rollout image from being reclaimed immediately while a rollout or short-lived workload is still progressing.

imageMaximumGCAge is a separate age-based limit. It allows an image to qualify for cleanup after remaining unused for the configured duration even if the disk usage high threshold has not been crossed. The documented default is 0s, which disables maximum-age cleanup. This setting does not expire an image that is actively being used by a container. It is not a substitute for checking whether a cold node can pull the image within the application’s recovery objective.

The maximum-age tracking has an operational caveat: a kubelet restart resets the tracked unused age. After a restart, the kubelet waits for the full configured maximum age before an image qualifies through that age-based condition. Do not treat imageMaximumGCAge as an exact deletion deadline or as a timestamp persisted independently of kubelet state.

Choose age values from observed rollout and recovery patterns. If a deployment uses a small set of large images and nodes are replaced frequently, retaining every unused image for a long period can consume meaningful disk. If rollout bursts depend on warm caches, an aggressive age or threshold policy can increase registry traffic and cold-start time. Measure the tradeoff using the node pool’s actual image sizes, pull throughput, registry availability, and deployment frequency.

Know which filesystem is under pressure

Image cleanup is easiest to reason about when the node’s storage layout is explicit. Kubernetes exposes filesystem identifiers used by kubelet accounting and eviction signals:

Identifier Typical contents Diagnostic question
nodefs Kubelet data, local ephemeral volumes, logs, and other node data Are logs, emptyDir, or local Pod data exhausting the node filesystem?
imagefs Container image read-only layers; it may also contain writable layers Is runtime image storage on a separate filesystem, and is it the constrained mount?
containerfs Container writable layers in supported split-image layouts Is writable container data separated from image layers, or does it share a mount?

Those names are logical identifiers as kubelet observes them. Two or all three can refer to the same physical filesystem; the labels do not prove that three separate disks exist. The container runtime configures where it stores image and writable layers. Moving runtime data is therefore a runtime and node-image change, not merely a kubelet YAML edit.

For the Kubernetes v1.37 documentation, use of containerfs requires the KubeletSeparateDiskGC feature gate, and CRI-O v1.29 or later is the documented runtime offering support for that filesystem signal. Check the exact Kubernetes and CRI runtime version, feature-gate state, and managed-platform support before designing around containerfs; do not generalize this combination to every runtime or an older release.

The filesystem layout changes which pressure signal rises first and which cleanup can reclaim it. If image layers live on their own filesystem, deleting images may relieve imagefs without freeing a full nodefs used by logs and local volumes. If everything shares one filesystem, image growth competes directly with those other consumers. A low image-cache reading also does not rule out inode exhaustion, log growth, writable layers, or a separate mount reaching its eviction threshold.

Distinguish image collection from Pod eviction and ephemeral-storage accounting

Image garbage collection removes eligible, unused image data. Node-pressure eviction selects Pods when configured or default node signals cross their thresholds; it is a separate mechanism with its own ranking and grace-period behavior. Local ephemeral-storage accounting tracks supported Pod-local data such as writable layers, container logs, and disk-backed emptyDir volumes. Those mechanisms can be related during a disk incident, but they do not share one cleanup policy.

Do not diagnose every DiskPressure condition by lowering image thresholds. First identify the affected signal and mount, then determine whether the constrained bytes or inodes belong to image layers, writable layers, logs, local volumes, or non-Kubernetes node processes. If the image filesystem is full, the kubelet may reclaim unused images; if the node filesystem is full of logs or emptyDir data, image reclamation on a separate mount may not help. If there are no unused images to delete, the image manager cannot manufacture free space.

Likewise, image garbage collection is not registry cleanup. Deleting a local layer from one node does not delete the registry tag or manifest; retention in the registry requires a separate, explicit registry policy. Prefer digest-pinned workload references where a stable artifact identity is required, and retain the corresponding artifact in a registry that meets rollback and disaster-recovery needs.

Observe safely before changing thresholds

Start with read-only observations from the control plane and the node’s supported management channel:

kubectl describe node worker-17
kubectl get node worker-17 -o jsonpath='{.status.conditions[?(@.type=="DiskPressure")].status}{"\n"}'
kubectl get pods -A --field-selector spec.nodeName=worker-17 -o wide

These commands help establish node condition and workload placement; they do not identify the exact filesystem mount or image directory. Use the operating system or managed-node diagnostics to map runtime storage paths to mounts and inspect byte and inode capacity. Follow the provider’s supported method to inspect kubelet configuration and logs. Avoid broad rm commands against runtime storage: manually deleting content can desynchronize the runtime’s metadata from files on disk.

For a controlled test, record the node’s filesystem layout, available bytes and inodes, image inventory, kubelet thresholds, Pod image digests, and pull duration. Trigger a representative rollout in a non-critical pool and check which images are retained, what gets pulled, whether the intended filesystem recovers, and whether startup/error budgets remain within target. A list of local images is useful inventory but not proof that a proposed deletion is safe or that a later pull will succeed.

Size the cache around recovery, not a percentage alone

Image cache policy should reflect how quickly replacement capacity must become ready. Estimate the image bytes likely to be needed by the expected burst of new or rescheduled Pods, then account for concurrent pulls, layer sharing, image size variability, and the free-space reserve required by logs and writable data. Shared layers reduce storage only when they are identical content and retained by the runtime; they do not remove the need to validate worst-case disk use.

Capacity tests should include a cold node, a registry slowdown, a failed pull followed by retry, and a rollout that temporarily references both old and new image digests. If the service cannot tolerate the cold-pull time, address the recovery path deliberately: improve registry availability and network capacity, stage images through the managed platform’s supported mechanism, or retain more cache where node capacity allows. Do not rely on external cron jobs that prune runtime storage behind kubelet’s back.

Define acceptance criteria before changing thresholds: no unexpected eviction, bounded node filesystem and inode use, known cold-start time, successful rollback to the prior digest, and no sustained registry saturation. Keep the configuration only if the measured disk/recovery tradeoff is acceptable for the workload and node pool.

Production checklist

  • Verify the Kubernetes release, kubelet configuration API version, runtime, and provider-supported configuration path.
  • Confirm high threshold is greater than low threshold and decide whether maximum unused age is actually needed.
  • Map nodefs, imagefs, and containerfs to real mounts; do not infer separate disks from the names.
  • Monitor bytes and inodes, image pull duration/failures, node condition, Pod evictions, and registry load together.
  • Test cold-node recovery and rollback with immutable image references in a non-critical pool.
  • Keep image cleanup under kubelet/runtime ownership; manage registry retention as a separate policy.
  • Roll out progressively and preserve enough evidence to revert the node-pool configuration safely.

Image garbage collection is healthy when it is predictable, observed, and boring: unused data is reclaimed within the designed storage budget, active workloads are left intact by the unused-image policy, and the next required image can still be obtained within the service’s recovery objective.

Related:

Sources:

Comments