Kubernetes Graceful Node Shutdown: Kubelet Budgets, Priority, and Limits
Configure kubelet graceful node shutdown with systemd constraints, Pod priorities and deadlines, then test failure boundaries without confusing it with drain.
Kubernetes graceful node shutdown is a kubelet response to an operating-system shutdown signal. It gives the node a bounded opportunity to terminate local Pods before the machine powers off. It is not a replacement for kubectl drain, a PodDisruptionBudget-aware maintenance workflow, or a cloud provider’s instance-termination coordination. Its success depends on the shutdown mechanism reaching kubelet in time, kubelet being configured with nonzero budgets, and the host actually granting the delay it requests.
The feature is easy to overestimate because a successful configuration test can look like a full node evacuation. In reality, kubelet is managing processes on one host under a deadline. It cannot manufacture replacement capacity, guarantee a workload’s own shutdown logic finishes, or intercept every abrupt power loss or provider action. Plan it as one layer in a node lifecycle strategy.
The shutdown signal and kubelet’s role
On Linux, graceful node shutdown uses systemd inhibitor locks. When systemd begins an orderly shutdown, kubelet can delay it, update the Node’s readiness state, and start terminating Pods running locally. During detected shutdown, kubelet rejects Pods at local admission, including Pods already bound to that node. This is distinct from a cluster drain: there may still be Pods pending placement, and scheduler behavior around tolerations should not be treated as a hard evacuation barrier. Keep capacity, taints, and the infrastructure termination workflow coordinated independently.
Linux support is Beta and enabled by default upstream since Kubernetes v1.21, but its timing settings default to zero, which means graceful termination is not active until configured. The systemd dependency matters: an API-triggered deletion, abrupt kill -9, kernel panic, power cut, or platform action that bypasses the relevant logind/shutdown path may not give kubelet the expected signal. A cloud VM’s termination deadline can also be shorter than the requested delay. Confirm the exact shutdown path used by the OS image and provider instead of assuming every reboot or instance stop is equivalent.
The ordinary graceful shutdown configuration has two time budgets:
- shutdownGracePeriod is the total time kubelet asks the system to delay shutdown for Pod termination.
- shutdownGracePeriodCriticalPods is the part reserved for critical Pods, and must be less than the total.
For example, shutdownGracePeriod: 90s and shutdownGracePeriodCriticalPods: 15s leave 75 seconds for regular Pods and the final 15 seconds for critical Pods. The periods are slices of a single total deadline, not 90 seconds plus 15 more seconds. Kubelet first terminates regular Pods and then critical Pods. When either value is zero by default, merely enabling the feature gate does not create time for the workload to shut down.
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
shutdownGracePeriod: 90s
shutdownGracePeriodCriticalPods: 15s
This is a kubelet configuration fragment for a Linux node, not a Pod manifest. Use the configuration delivery mechanism supported by the distribution. Managed Kubernetes providers may render kubelet configuration from a node group or launch template; editing a host file directly may be ignored or overwritten.
Make priority represent recovery order
The basic two-stage behavior distinguishes regular Pods from critical Pods. Kubernetes reserves the system-cluster-critical and system-node-critical PriorityClasses for cluster-critical workloads. Do not label an application critical just to win a shutdown race: PriorityClass can also affect scheduling and preemption, and overuse weakens the meaning of the priority system. Decide which services must remain alive longest, which can restart safely elsewhere, and which are required for node or cluster operation.
For finer-grained shutdown periods, kubelet supports shutdownGracePeriodByPodPriority. Each configured threshold applies a per-Pod grace period to a priority range. With descending thresholds of 100000, 10000, and 0, Pods at priority 100000 or higher receive the first listed budget; those from 10000 up to but excluding 100000 receive the second; and lower priorities receive the final budget. If a priority range has no local Pods, kubelet skips that empty range.
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
shutdownGracePeriodByPodPriority:
- priority: 100000
shutdownGracePeriodSeconds: 20
- priority: 10000
shutdownGracePeriodSeconds: 45
- priority: 0
shutdownGracePeriodSeconds: 75
Use this list instead of the two standard shutdown period fields; kubelet configuration does not allow both mechanisms to be set together. The priority-based feature requires GracefulNodeShutdownBasedOnPodPriority. Upstream documents it as Beta since v1.24 and enabled by default in v1.37; check the exact release and provider build before relying on that state. If the feature is enabled without a priority budget list, it performs no priority ordering action.
These are not additive serial windows. Kubelet initiates shutdown on local Pods with grace periods derived from their priority ranges, then waits for them to exit; the shutdown-inhibit time is bounded by the largest configured period represented by local Pods, not the sum of every list entry. Set the whole budget so it fits within systemd’s effective inhibitor delay and the external infrastructure deadline. A custom list cannot buy time the host or cloud platform will not grant.
Respect the Pod-level deadline and application behavior
Kubelet follows the normal Pod termination process while shutting down a node, including the Pod’s configured grace period and container lifecycle handling. This means a long preStop hook or an application that ignores termination can consume the budget. When the node-wide window expires, kubelet and the operating system cannot promise the process gets an unbounded extension. Align the node budget with application shutdown time, but remember that every local Pod competes inside the host’s finite overall window.
A Pod’s terminationGracePeriodSeconds is not a guarantee that it receives that full duration during a node shutdown. The node-level deadline is a separate cap. Nor does graceful node shutdown guarantee in-flight requests have drained from every load balancer, that a volume detach completed, or that a replacement Pod is already ready. Applications still need idempotent recovery, durable state, correct readiness behavior, and a tested termination handler. For planned maintenance, cordon and drain the node first so the control plane can schedule replacements and the eviction API can apply disruption policy before the host shutdown begins.
Review interactions with OS maintenance tools. Kubernetes documents a Debian unattended-upgrades configuration that can conflict with the systemd inhibitor lock when the requested shutdown grace period exceeds 30 seconds. A kubelet log claiming successful startup is not proof the machine manager will honor the full interval. Verify systemd/logind configuration, service permissions, provider termination notices, and the final poweroff deadline on the actual node image.
Observe what happened, not just what was configured
Before testing, capture the kubelet config source and effective version, feature-gate state, node Ready condition, Pod list and priorities, application shutdown logs, events, and system journal timestamps. During a test, observe:
- When the OS begins an orderly shutdown and when kubelet detects it.
- When node readiness changes and whether new Pods are rejected at the node.
- Which priority bands terminate first and their actual exit times.
- Whether preStop and application cleanup finish before their deadlines.
- When the shutdown lock is released and when the host actually powers off.
- Whether controllers create replacements and whether those replacements become Ready elsewhere.
The kubelet exposes graceful_shutdown_start_time_seconds and graceful_shutdown_end_time_seconds metrics for priority-based shutdown monitoring in current upstream documentation. Confirm these metrics exist in your kubelet version and scrape configuration before using them in an alert. Correlate metrics with systemd journal messages and API events; no single signal proves the full shutdown sequence succeeded.
Run acceptance tests on a disposable node pool. Test the exact orderly shutdown command used by the operating system, the provider’s normal maintenance/termination workflow, a cancellation if the environment supports one, and an abrupt shutdown to confirm the expected boundary. Verify that a cancelled shutdown does not resurrect Pods that already began termination; replacements may still be needed. Test with representative critical and regular priorities, a slow shutdown hook, a terminating application, and constrained replacement capacity. Do not use production nodes as the first place to discover a logind or provider timeout.
Know when this feature cannot help
Graceful node shutdown cannot reliably protect workloads from sudden power loss, host crash, abrupt termination outside the inhibitor path, kubelet being unable to contact the API server, or a host shutdown deadline shorter than the configured grace. Non-graceful node shutdown handling addresses a different failure mode: it helps the control plane recover Pods from a node that shut down without kubelet’s graceful path, but it does not give processes extra time to flush application data. Storage attach/detach, StatefulSet identity, and fencing still require their own recovery tests.
Likewise, local kubelet shutdown is not equivalent to a voluntary API eviction. Do not promise that a PodDisruptionBudget will reserve enough replacement capacity or prevent a host from powering off. For operator-initiated maintenance, the safer order is to cordon, drain with the appropriate PDB and timeout policy, confirm workload health elsewhere, and only then shut down the machine. For cloud-driven termination, use provider-supported lifecycle hooks or node termination handling plus sufficient spare capacity; treat kubelet’s local grace window as the final host-side opportunity, not the only one.
Record the budget, supported shutdown paths, provider deadline, Pod priority mapping, measured termination duration, and recovery owner in the node-pool runbook. Re-test after kubelet, OS, systemd, container runtime, or provider lifecycle changes. A graceful shutdown feature is production-ready only when the entire path from shutdown signal to service recovery has been observed under its real external deadline.
Related:
- Kubernetes Pod Termination: Graceful Shutdown and Endpoint Draining
- Karpenter Node Lifecycle: NodePools, NodeClaims, Scheduling, Consolidation, and Disruption
Sources: