Kubernetes Jobs: Completion, Retry, and Pod Failure Policy
Design reliable Kubernetes Jobs with explicit completion targets, bounded retries, indexed work, failure classification, deadlines, and cleanup.
A Kubernetes Job represents work that should run to completion rather than remain available as a continuously reconciled service. The Job controller creates Pods until the requested work succeeds or reaches a failure policy. That controller-level guarantee is useful, but it is not an exactly-once execution guarantee for an external side effect. A Pod can be replaced after disruption, a process can complete an external write before its success is recorded, and a later retry can execute the same logical task again.
Production Job design therefore starts with the workload contract: how many successful completions are required, how much parallelism is safe, which failures are retryable, how long the work may run, and how its output is made idempotent or deduplicated. The YAML is only one part of that contract. Durable progress checkpoints, unique operation keys, and observable exit reasons often matter more than a high retry count.
Choose completion and parallelism semantics
For a single bounded task, completions: 1 and parallelism: 1 are usually a clear starting point. For a finite set of independent items, use a larger completion count and decide whether workers can process any item or need a stable index. completionMode: Indexed assigns each completion an index that the process can use to select a deterministic shard or partition. The index is not a substitute for a durable work ledger: Pods can retry, and application-side writes still need to tolerate replay.
parallelism is a concurrency ceiling, not a promise that exactly that many Pods will run at every instant. Capacity, quotas, scheduling, image pulls, and API availability all affect actual concurrency. A worker system should be safe when fewer Pods start than requested and should avoid assuming that a specific worker owns a partition forever. For shared queues, lease or claim work with expiration and fencing so that a delayed worker cannot overwrite a newer owner.
apiVersion: batch/v1
kind: Job
metadata:
name: invoice-reconciliation
namespace: batch
spec:
completions: 12
parallelism: 3
completionMode: Indexed
backoffLimitPerIndex: 2
maxFailedIndexes: 0
activeDeadlineSeconds: 7200
ttlSecondsAfterFinished: 86400
template:
spec:
restartPolicy: Never
containers:
- name: reconcile
image: registry.example.invalid/billing/reconcile@sha256:REPLACE_WITH_DIGEST
args: ["--index=$(JOB_COMPLETION_INDEX)"]
This manifest is a template, not a deploy-ready image reference. Replace the example digest, align deadline and retention with the actual workload, and verify the Kubernetes version supports every field you use. For an Indexed Job, Kubernetes makes the completion index available in JOB_COMPLETION_INDEX; the example passes it to a hypothetical application-specific flag using Kubernetes command/argument expansion. The Indexed mode gives each completion an index, while backoffLimitPerIndex limits retries for each index independently. Setting maxFailedIndexes: 0 states that no failed shard is acceptable. If partial completion is useful, choose a deliberate nonzero allowance and make downstream consumers understand incomplete output.
The Job documentation warns that even a Job with one completion, one parallel Pod, and restartPolicy: Never can sometimes start the same program twice. Treat any interaction with a database, object store, payment system, or queue as potentially repeated. Use an idempotency key derived from the logical work item, commit output atomically where possible, and record completion so a retry can safely detect prior success.
Bound retries and total runtime
backoffLimit bounds failed Pod retries for a Job. Its default is six unless per-index retry behavior changes the default. The controller increases the delay between retries exponentially and caps the delay, but repeated retries still consume cluster capacity and may repeatedly hit a permanent application error. Choose the retry limit from the failure modes you can recover from, not as a substitute for a timeout or alert.
activeDeadlineSeconds bounds total active execution time for the Job. Once that limit is reached, the Job stops creating more Pods even if the failure count has not reached the retry limit. Use both controls when there is a meaningful wall-clock deadline and a separate failure budget. A program that can hang indefinitely needs its own timeout for network calls and child processes; controller deadline enforcement is not a substitute for cooperative shutdown inside the container.
Avoid timeouts that are shorter than expected scheduling and startup delay. A Pod can spend time Pending because of a quota, unavailable node capacity, image pull delay, or volume attach. If the Job is a daily accounting task, measure end-to-end duration from creation through completion and size deadlines to the operational window. Alert before the deadline if a task is making no progress, and capture the Pod events and application logs before cleanup removes them.
Classify failures only when the evidence supports it
The default retry behavior counts failed Pods toward the Job’s backoff budget. A podFailurePolicy can distinguish failure classes using exit codes or Pod conditions. Its rules are evaluated in order; the first matching rule determines the action. FailJob terminates the Job for an unrecoverable condition, Ignore does not increment the retry counter and requests replacement work, Count uses normal failure accounting, and FailIndex can mark an index failed when per-index retry behavior is enabled.
Pod failure policy requires the Pod template to use restartPolicy: Never. Treat rule ordering as executable policy: a broad rule placed before a specific one can consume the failure and prevent the later rule from matching. Name the container in an exit-code rule where possible, distinguish application exit codes from infrastructure conditions, and test what happens when a Pod is deleted or evicted. Do not classify all nonzero exits as permanent; a temporary dependency failure and invalid input need different responses.
podFailurePolicy:
rules:
- action: Ignore
onPodConditions:
- type: DisruptionTarget
- action: FailJob
onExitCodes:
containerName: reconcile
operator: In
values: [64, 65]
The snippet belongs under spec and illustrates two classifications: disruption is ignored for retry accounting, while selected application exit codes fail the overall Job. Use exit-code meanings that the application actually documents. If the process returns the same code for a transient error and invalid input, the controller cannot infer the distinction safely. Test terminating Pods as well as already-terminal failures, because controller accounting and replacement timing are tied to Pod lifecycle state.
When podFailurePolicy is in use, the controller evaluates failed Pods based on their terminal state. Starting with Kubernetes v1.28, the Job controller creates a replacement as soon as termination of the old Pod becomes apparent, but does not count that Pod as a failure or evaluate the failure policy until the old Pod reaches a terminal phase. The old process can therefore still be shutting down while replacement work starts; retain idempotency and concurrency protections. Verify the exact behavior on the cluster version you operate, especially during upgrades and when using a managed distribution with its own controller lifecycle.
Preserve progress and make output recoverable
A Job object records controller status, not a complete application checkpoint. For long-running work, persist checkpoints outside the Pod and include a monotonically increasing generation or fencing token. If a retry begins after the previous worker performed an external action but before its status was recorded, the action must be safely repeatable or detectable by a unique request ID.
For indexed Jobs, persist results by completion index and validate that every required index produced one successful result. Do not infer completeness from “the Job is no longer active” alone: inspect its Complete or Failed condition and compare succeeded and failed indexes to the expected range. For non-indexed work from a queue, measure claimed, acknowledged, retried, and dead-lettered items separately. A green Job status cannot prove downstream data is correct.
Keep worker output machine-readable and include Job name, namespace, UID, completion index when present, input range, attempt number, and application version. Avoid relying on a Pod name as a durable work identity because a retry creates another Pod. External operations should use a stable logical key and record the producing Job UID or business batch ID as provenance.
Retain diagnostics, then clean up deliberately
ttlSecondsAfterFinished makes a finished Job eligible for automatic deletion after a delay. It helps prevent abandoned Jobs and their Pods from accumulating, but aggressive cleanup can remove logs or events before an incident is investigated. Retain centralized logs and metrics outside the Job objects, choose a TTL that fits the diagnostic window, and ensure failed Jobs alert before they are removed.
CronJob-created Jobs have a different retention context: the CronJob’s successful and failed history limits manage how many completed Jobs remain. Avoid overlapping cleanup policies that make it unclear which controller removes evidence. For manually created batch Jobs, set a TTL or operate a cleanup controller and document whether the Job, Pods, logs, and external results have separate retention periods.
Use kubectl get jobs, kubectl describe job, and kubectl get pods -l job-name=... to inspect progress. Capture Pod status, exit codes, events, logs from every failed attempt, and relevant application checkpoints before deleting a failed run. A retry should be a deliberate recovery action, not a loop that erases the first useful diagnostic.
Validate the failure matrix before production
In a staging cluster, test success, transient dependency failure, permanent input rejection, Pod eviction, unschedulable capacity, image pull failure, deadline expiration, and controller restart while a Pod is terminating. Verify the retry counters, podFailurePolicy ordering, Job conditions, failed-index status, cleanup delay, and alert path. Confirm that re-running the same business batch does not duplicate side effects.
Treat the Job as a bounded orchestrator around an idempotent program. Explicit completion semantics, a retry budget, a runtime deadline, durable progress, and preserved diagnostics turn a batch controller from “keep launching Pods” into an operationally understandable contract.
Related:
- Kubernetes CronJob Execution Semantics: Designing Reliable Scheduled Work
- Kubernetes Pod Termination: Graceful Shutdown and Endpoint Draining
Sources: