Kubernetes Node-Pressure Eviction: Thresholds, Ranking, and Recovery
Diagnose Kubernetes node-pressure evictions, configure kubelet thresholds safely, and distinguish eviction from OOM kills, preemption, and API disruptions.
Kubernetes node-pressure eviction is a kubelet recovery mechanism: when a node runs short of a monitored resource, the kubelet first tries node-level reclamation and then terminates Pods to protect the node. It is not the scheduler choosing a victim for a new Pod, and it is not the same operation as an administrator calling the Eviction API. Those distinctions matter during an incident because a Failed Pod, an OOMKilled container, a MemoryPressure condition, and a preempted workload have different causes and recovery paths.
This guide focuses on the node-local path: what the kubelet measures, how thresholds work, how it orders victims, what protection mechanisms do not cover, and how to design a response that preserves evidence while restoring capacity. Kubelet behavior and available signals evolve; validate configuration against the Kubernetes release and node operating system actually in use.
Identify which mechanism acted
Start by separating four events that are often reported simply as “the Pod was evicted.”
| Mechanism | Decision maker | Typical evidence | What it does |
|---|---|---|---|
| Node-pressure eviction | kubelet on the affected node | Pod phase Failed, reason Evicted, node pressure condition or kubelet event |
Terminates a Pod to reclaim node memory, filesystem space/inodes, or PIDs |
| Container OOM kill | Linux kernel or container memory controller | Container last state Terminated, reason OOMKilled; restart may follow policy |
Kills a container process; the Pod can remain active and the kubelet may restart the container |
| Scheduler preemption | scheduler | Pending high-priority Pod events and lower-priority victims | Removes lower-priority Pods to make room for a higher-priority Pod on a node |
| API-initiated eviction | API server through the Eviction subresource | Eviction request result, workload-controller replacement | Performs a voluntary disruption request and can honor a PodDisruptionBudget |
Node-pressure eviction sets the selected Pod’s phase to Failed and terminates it. The kubelet does not respect a PodDisruptionBudget for this involuntary node-local action. A PDB therefore cannot be treated as a shield against memory, disk, inode, or PID starvation. Soft and hard kubelet thresholds also differ in termination behavior: hard thresholds use a zero-second grace period; soft thresholds use a grace period capped by eviction-max-pod-grace-period when configured. Do not assume an application’s normal shutdown window will always be available.
Signals, thresholds, and node conditions
The kubelet compares eviction signals with configured thresholds. Common signals include memory.available, nodefs.available, nodefs.inodesFree, imagefs.available, imagefs.inodesFree, containerfs.available, containerfs.inodesFree, and pid.available. Not every filesystem signal applies to every node layout or platform. For example, Linux inode signals do not describe Windows nodes, and the kubelet discovers only supported filesystem arrangements from the node and container runtime.
For Linux memory eviction, memory.available is calculated from cgroup-aware node statistics rather than the output of free -m inside an arbitrary container. The kubelet also treats some file-backed memory as reclaimable. This is one reason a host-level memory dashboard or a container’s RSS alone may not reproduce the kubelet’s decision exactly. Compare the node condition, kubelet events and logs, and the metrics source used by the kubelet before changing a threshold.
Thresholds can be absolute quantities or percentages. A soft threshold requires a matching grace period; the condition must persist through that period before eviction begins. A hard threshold has no grace period. The default thresholds are release- and platform-specific. In current upstream documentation, the Linux and Windows memory defaults differ, and filesystem defaults are also listed. If you override one threshold, do not assume every other default remains active: depending on kubelet configuration, unspecified thresholds can become zero. Supply the complete intended set or deliberately enable the documented merge-default behavior, then inspect the effective configuration on the actual nodes.
The kubelet reports MemoryPressure, DiskPressure, or PIDPressure when the corresponding signal reaches a threshold, and the control plane maps those conditions to taints. These conditions are valuable incident context, but they are not proof that the affected application caused the pressure. System daemons, logs, image layers, local volumes, inode exhaustion, or workloads that exceed requests can all contribute.
What the kubelet tries before evicting Pods
For disk pressure, the kubelet attempts to reclaim node-level resources first. Depending on whether node, image, and container filesystems share or split storage, it can garbage-collect dead containers and Pods, then delete unused images. It does not arbitrarily clean every mounted disk. A node with a separate application data mount can still fill that mount without the kubelet reclaiming it.
Only after node-level reclamation fails to bring a signal back below its threshold does the kubelet start evicting end-user Pods. A Pod’s local volumes, logs, and writable layers may count differently depending on the filesystem layout. ephemeral-storage requests and limits help account for local ephemeral use, but they do not make every node filesystem or external volume safe. Monitor the resource that is actually exhausted and map it to its backing filesystem.
For memory, the node may reach a kernel OOM condition before the kubelet observes or acts on pressure. A container-level memory limit can trigger a cgroup OOM kill while the node still has capacity; node-level pressure can instead cause the kubelet to fail a Pod. These can happen close together during a rapid spike, so preserve termination state and events rather than diagnosing solely from a dashboard snapshot.
Eviction order is not simply QoS order
For resource signals that have requests, kubelet ranking considers whether a Pod’s usage exceeds its requests, then Pod priority, then how far usage is above requests. As a broad result, BestEffort and Burstable Pods using more than their requests are considered before Guaranteed Pods and Burstable Pods staying below their requests. Within those broad groups, priority and relative overage matter.
Although QoS class can help estimate memory behavior, the kubelet does not use the class label itself as the ranking algorithm. This is especially important for ephemeral storage: QoS classes are not based on ephemeral-storage requests, so “Guaranteed” does not mean immune to disk-pressure eviction. Inode and PID starvation have no corresponding per-Pod requests, so relative Pod priority is used for their eviction ordering.
Priority is not an absolute guarantee either. If system daemons consume more than their reserved capacity and the remaining Pods are below requests, the kubelet can still need to evict one to protect node stability; it selects lower-priority Pods first in that situation. Avoid assigning critical priority classes indiscriminately. A cluster full of “critical” workloads has merely hidden the trade-off, not eliminated the resource shortage.
A measured triage sequence
Capture state before deleting or restarting the affected objects. Start with the Pod’s status, last container state, events, node, and node conditions:
kubectl get pod -n production api-7f9d8b6d9c-abcde -o wide
kubectl describe pod -n production api-7f9d8b6d9c-abcde
kubectl get pod -n production api-7f9d8b6d9c-abcde \
-o jsonpath='{.status.phase}{"\t"}{.status.reason}{"\t"}{.status.message}{"\n"}'
kubectl describe node worker-07
kubectl get events -A --sort-by=.lastTimestamp
Then compare the declared requests and limits with observed use and node allocatable capacity. Check namespace-wide requests as well as the victim: a single Pod may be selected because it is over its request, while many other Pods collectively create the pressure. For disk, inspect node filesystem layout and ephemeral usage; a generic “disk full” alert does not tell you whether nodefs, imagefs, containerfs, an inode pool, or an unrelated mount is the constrained resource. For memory, correlate kubelet and kernel evidence with container termination state to distinguish eviction from a cgroup OOM kill.
Useful evidence includes the node condition transition, kubelet eviction messages, Pod events, resource metrics around the event, and the exact deployed kubelet configuration. Managed Kubernetes providers may expose only part of the host evidence; use their node diagnostics rather than assuming a missing local log means the event did not occur. Record the Kubernetes and kubelet versions because defaults and supported filesystem behavior can change between releases.
Configure thresholds as an operating policy
Treat eviction thresholds as a node reliability budget, not as a magic workload limit. A memory threshold should leave enough room for the kernel, kubelet, runtime, system daemons, and short-lived bursts. Reserve those resources with the appropriate node configuration and ensure scheduler-allocatable capacity reflects what workloads may safely consume. Otherwise the scheduler can place Pods whose requests fit on paper but whose combined real use causes immediate eviction.
Soft thresholds are useful when early warning and bounded graceful shutdown are operationally meaningful. Set the signal, grace period, and maximum grace deliberately; test the exact termination sequence and account for the preStop hook and application shutdown time. Hard thresholds are an emergency boundary, not a graceful-drain feature. Ensure an alert fires before either threshold, and test it on the same node image and kubelet release used in production.
For disk and inode pressure, capacity planning must include image churn, container logs, writable layers, and emptyDir or other local ephemeral use. Configure workload ephemeral-storage requests and limits where appropriate, rotate and retain logs intentionally, and alert on both free bytes and free inodes. A node can have gigabytes free but still fail to create files when its inode pool is exhausted.
A safe configuration review checklist
- Confirm which signal and filesystem correspond to the reported pressure.
- Read the effective kubelet configuration, not only a Helm value or bootstrap template.
- If overriding defaults, verify every threshold that should remain active or explicitly use the supported merge behavior.
- Keep
system-reserved,kube-reserved, scheduler allocatable capacity, and eviction thresholds consistent. - Test soft and hard behavior in a disposable node pool, including application shutdown and workload replacement.
- Alert on pressure conditions, threshold proximity, eviction counts, and resource use above requests.
- Review provider-specific node upgrades because kubelet defaults and storage discovery can change.
Recovery without masking the cause
Controllers such as Deployments and StatefulSets can create replacement Pods after a Pod fails, but a replacement on the same pressured node may fail again. Check placement and node health before repeatedly deleting the symptom. If a node is persistently unhealthy, cordon or drain it according to the platform’s supported procedure, preserve diagnostic data, and restore capacity before rescheduling the workload. A voluntary drain can interact with PDBs; node-pressure eviction does not, so they must not be conflated.
Fix the exhausted budget: right-size requests from measured use, set realistic limits where they are effective, reduce or isolate noisy workloads, add node capacity, reserve system resources, or correct disk/log lifecycle. Raising an eviction threshold without changing the underlying budget can simply delay the same failure until the node has less recovery room. Lowering it may increase evictions. Review the change as a capacity and availability decision, then monitor both the pressure signal and the application’s service-level indicators.
The durable mental model is simple: the kubelet protects a node locally, the scheduler places Pods using declared requests and node allocatable capacity, and the kernel can still kill a container before kubelet eviction completes. PDBs govern voluntary API disruptions, not emergency node-pressure eviction. Identify which boundary acted first, preserve the evidence, and change the budget or workload behavior that caused the signal rather than treating every termination as the same Kubernetes event.
Related:
- Kubernetes QoS Classes: How Requests and Limits Shape Eviction and Scheduling
- Kubernetes Pod Priority and Preemption: Scheduling Order, Victims, and PDB Limits
Sources: