Kubernetes Probe Semantics: Startup, Liveness, and Readiness
Configure Kubernetes probes around initialization, process health, and traffic eligibility without causing restart loops or false outages.
Kubernetes probes answer three different questions about a container: has it finished starting, should it be restarted, and should it receive traffic? Treating these as interchangeable “health checks” is a common cause of avoidable outages. A readiness failure removes a Pod from Service traffic while leaving the container running; a liveness or startup failure can cause the kubelet to terminate and restart the container.
Probe design is an application contract. The kubelet can observe an HTTP response, an open TCP port, a gRPC health result, or the exit status of an exec command. It cannot infer from a successful socket connection that a database transaction will succeed, nor can it know whether restarting an overloaded service will make the system healthier.
Startup protects slow initialization
A startup probe gives a slow-initializing process a separate allowance to become operational. While that probe has not succeeded, the kubelet does not run that container’s liveness or readiness probes. This avoids forcing one very long liveness delay on every future restart just because the first initialization can take time.
Choose a startup failure budget that matches measured cold-start behavior, not the slowest imaginable outage. If initialization never completes, the failed startup probe still leads to container termination and the Pod’s restart policy. Keep logs and startup progress visible so repeated startup failure is diagnosable rather than a silent cycle.
Liveness is for unrecoverable lack of progress
A liveness probe should detect a process that is running but cannot make progress in a way that restarting can repair. A local deadlock or an internal event loop that has stopped advancing may be appropriate signals. A dependency outage, full work queue, or temporary database failure usually is not: restarting every replica can multiply load on the same failing dependency.
Liveness failures can become cascading failures under load. If a service slows because it has too little capacity, an aggressive timeout may repeatedly restart it, reducing the capacity that remains. Measure the health handler itself, size its timeout, and test under representative saturation. Avoid making the liveness check perform a deep request through every downstream service.
Readiness controls traffic, not process lifetime
Readiness asks whether the instance should receive normal Service traffic now. A process can be alive but not ready while warming caches, loading a model, restoring local state, or temporarily shedding load. A failed readiness result marks the Pod not ready and removes it from matching Service EndpointSlices; it does not by itself restart the container.
Do not use readiness to hide permanent configuration errors indefinitely. Emit a precise reason, alert on prolonged unready time, and decide whether the right response is to repair configuration, roll back, or stop the rollout. Also keep the readiness handler inexpensive and bounded so a growing process leak in an exec-based check does not starve the application.
Example with a bounded startup window
This manifest uses an HTTP startup probe for initialization, a shallow liveness endpoint for process progress, and a separate readiness endpoint for traffic eligibility. The image reference and paths are placeholders; replace them with an image that implements the probe semantics before applying the manifest.
apiVersion: v1
kind: Pod
metadata:
name: api
spec:
containers:
- name: api
image: example.invalid/api:reviewed
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /startup
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 36
livenessProbe:
httpGet:
path: /live
port: http
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2
The startup settings permit roughly three minutes of unsuccessful startup checks before failure, subject to kubelet timing. The liveness and readiness endpoints should be designed separately even if they share internal code. Liveness should not require a remote database to be reachable; readiness may include a bounded dependency check only if that dependency truly determines whether this replica can serve useful traffic.
Probe-level terminationGracePeriodSeconds was introduced in Kubernetes 1.25 and is stable from 1.28. It is supported for startup and liveness probes, but not for readiness probes. Use it only when a failed container needs a distinct shutdown grace period and verify that the cluster API version accepts the field. Do not copy a newer field into an older cluster without checking its schema.
Choose a mechanism deliberately
HTTP checks are easy to inspect but can accidentally test a reverse proxy or a cached response instead of application state. TCP checks establish only that a port accepts a connection; they do not prove the service can parse a request. gRPC checks use the standard health protocol and expect SERVING. Exec checks can express a local condition, but the kubelet creates a process for each execution; dense clusters with frequent checks can pay measurable process and CPU overhead.
After deployment, inspect Pod events, probe results, restart counts, and EndpointSlice membership. Test startup with deliberately slow initialization, liveness with a controlled deadlock in a non-production environment, and readiness with recoverable overload or dependency failure. The expected outcomes must differ: startup eventually succeeds, liveness restarts only when recovery by restart is intended, and readiness removes traffic without restarting the process.
Related:
- Diagnosing and Fixing CrashLoopBackOff in Kubernetes
- How to Configure Pod Disruption Budgets in Kubernetes
Sources: