Kubernetes Pod Termination: Graceful Shutdown and Endpoint Draining
Trace the Pod deletion path through preStop, SIGTERM, EndpointSlice conditions, connection draining, and the final grace-period deadline.
A Kubernetes Pod deletion is a coordinated shutdown, not an instantaneous process kill. The API records the deletion and grace period; the kubelet on the node then runs the container shutdown path. At the same time, the control plane updates the Pod’s EndpointSlice state so traffic systems can stop selecting it for new ordinary requests. Those updates are asynchronous, so applications and proxies still need a deliberate draining strategy.
If a container has a preStop hook and the grace period is non-zero, the kubelet runs that hook first. The grace-period countdown has already started, so a slow hook consumes time the process would otherwise have for shutdown. After the hook completes, the runtime sends the stop signal, normally SIGTERM; after the deadline the process can be forcibly killed. Do not budget the hook and the application shutdown as if each receives the full configured period.
Let the application stop accepting new work
The application should handle termination by stopping admission of new requests, finishing or safely cancelling in-flight work, flushing durable state, and exiting before the deadline. The EndpointSlice marks a terminating endpoint as not ready for regular traffic; its serving condition can still communicate whether it is serving during termination for consumers that implement draining behavior.
spec:
terminationGracePeriodSeconds: 45
containers:
- name: api
image: example/api:1.4.2
lifecycle:
preStop:
httpGet:
path: /internal/drain
port: 8080
This is an illustrative fragment, not a universal hook design. A local drain endpoint must be authenticated or otherwise inaccessible to untrusted clients, idempotent, fast, and coordinated with the application’s signal handler. A fixed sleep is not proof that load balancers, service meshes, or external clients have stopped sending traffic.
Test the actual termination path
Delete a test Pod during sustained requests and observe the Pod, EndpointSlice conditions, application logs, proxy behavior, open connections, and termination timestamps. Verify that replacement capacity becomes ready before the old instance exits when availability requires overlap. Repeat with a slow hook and an application that exceeds its shutdown budget so forced termination is visible and monitored.
PodDisruptionBudgets govern eligible voluntary disruptions; they do not make every deletion graceful or guarantee a minimum number of running Pods during involuntary failures. Document which controller performs a rollout, eviction, node drain, or direct deletion, and test each mechanism that matters to the service.
Separate endpoint removal from process shutdown
Pod deletion has two related but independent effects. The API object receives a deletion timestamp and grace deadline; the kubelet observes that state on the node and runs the container shutdown path. Meanwhile, control-plane controllers update EndpointSlices. Terminating endpoints are retained long enough to expose their state, but ready is set to false for backward compatibility. A consumer that needs to drain existing sessions can inspect the serving condition, but only if that consumer implements the behavior. An application must not assume that every load balancer, proxy, DNS cache, or client sees the same update at the same instant.
That distinction explains a common rollout race: the process starts shutting down while some existing connections still arrive. Handle it in the application protocol. Stop accepting new work when shutdown begins, return an appropriate retry response where safe, drain established connections up to a bounded deadline, and make any in-flight operation either complete or fail in a recoverable way. Readiness probes help remove an unhealthy Pod from normal service selection, but they are not a substitute for a signal handler or a connection-draining policy.
Budget the complete grace period
The grace timer includes the preStop hook. If a hook is still running when the configured period expires, the kubelet may request a one-off two-second extension; that is a narrow safeguard, not extra capacity to design around. Once the hook completes, the runtime requests a stop for the container’s main process. SIGTERM is the usual default for common runtimes, but an image can define a STOPSIGNAL, and Kubernetes has version- and feature-gate-specific custom stop-signal support. Check the image and target cluster instead of treating one signal as universal.
In a multi-container Pod, do not infer an ordering between ordinary application containers: their stop requests can be processed asynchronously. If one process must remain alive while another exits, use a supported lifecycle relationship such as a native sidecar where appropriate, or coordinate explicitly in the application. A long hook in one container also consumes Pod termination time and can leave little budget for the rest of the shutdown path.
Choose the grace value from measured behavior: worst-case request duration, queue-drain time, durable flush time, and any hook work must all fit with margin. Increasing the value indefinitely is not a fix; slow termination can delay rollouts and node maintenance. Conversely, too small a value turns routine deployments into forced kills. Alert on Pods that repeatedly reach the deadline, and distinguish an application that ignored its signal from one that was still performing legitimate bounded cleanup.
Exercise the path you actually use
In a disposable environment, watch the object and endpoints while sending sustained traffic:
kubectl get pod -w -l app=api
kubectl get endpointslice -w -l kubernetes.io/service-name=api
kubectl delete pod api-0 --wait=false
The selector and Pod name are examples; use the workload’s real labels and test replica. Correlate the deletion timestamp, EndpointSlice terminating/ready/serving conditions, hook logs, application signal logs, proxy connection state, and final container exit code. Repeat through the production-relevant controller path: a Deployment rollout, an eviction or node drain, and a direct deletion do not exercise identical policy. A PodDisruptionBudget constrains qualifying eviction requests; it does not block every direct API deletion, protect against node loss, or guarantee that an application will finish before its grace deadline.
Use kubectl describe pod and the application’s structured shutdown logs to diagnose a timeout. If the endpoint disappears promptly but requests still fail, inspect the proxy and client retry behavior. If the endpoint stays eligible, inspect readiness and EndpointSlice conditions rather than adding an arbitrary sleep to the hook. If the container is killed, compare its stop time with the grace deadline and budget each stage independently. These observations turn shutdown from a best-effort guess into a testable availability contract.
Related:
- Kubernetes Probe Semantics: Startup, Liveness, and Readiness
- How to Configure Pod Disruption Budgets in Kubernetes
Sources: