Kubernetes CronJob Execution Semantics: Designing Reliable Scheduled Work
Design reliable Kubernetes scheduled work around duplicate starts, retries, concurrency, time zones, and observable recovery behavior.
A Kubernetes CronJob is a controller that creates Job objects on a schedule; each Job then manages Pods that run a task to completion. This separation matters operationally: the CronJob decides when to request work, while Job retry behavior governs attempts after a Job exists. Neither controller can make an external side effect exactly once. Kubernetes documents CronJob scheduling as approximate: under some conditions a schedule can create two Jobs or none. Treat each invocation as a request to process work safely, not as a unique business event.
Make retries and duplicate starts harmless
Job Pods can be replaced after node failure, eviction, or other interruption. Kubernetes explicitly warns that the same program can sometimes start twice even with one completion, one parallel Pod, and restartPolicy: Never. The Job controller also backs off and retries failed Pods up to its configured backoffLimit; that is an execution policy, not a transaction spanning your database or API.
Design the operation around a durable idempotency key such as report:2026-10-03T12:00Z, not a random identifier generated at process startup. Persist a unique constraint or claim for that key, and make the state transition and output commit atomic where possible. For external effects, use the destination’s idempotency mechanism or a transactional outbox. If the process crashes after sending a request but before recording success, a retry must be able to determine whether to repeat, reconcile, or safely skip the effect. Temporary files and partial output also need cleanup or resumable checkpoints.
Define overlap, lateness, and recovery explicitly
concurrencyPolicy applies only to Jobs created by that one CronJob. Allow permits overlap; Forbid skips a scheduled run while an earlier Job from that CronJob is active; Replace replaces that CronJob’s active Job with a new one. These are not global locks: a manually created Job or a second CronJob can still run the same application at the same time. Use application-level coordination when work must be exclusive across those paths.
Set startingDeadlineSeconds according to the business value of late work. After the deadline, that occurrence is skipped while future schedule times remain eligible. Leaving the field unset means there is no deadline, not that every missed occurrence will be replayed. The controller limits catch-up after extended outages; when it counts more than 100 missed schedules, it does not start a catch-up Job. With a deadline, it counts missed occurrences inside the recent deadline window rather than across the full interval since the last scheduled run, which can prevent a long outage from being treated as an unbounded backlog. Values below 10 seconds are unsafe: the controller checks schedules every 10 seconds, so such a CronJob may not be scheduled. Forbid can itself make schedule times count as missed while a previous run remains active. A deadline is therefore a freshness policy, not a durable queue.
Suspending a CronJob stops it from creating new Jobs but does not stop Jobs that are already running. Missed executions count as missed Jobs; when the CronJob is unsuspended with no starting deadline, those missed Jobs can be scheduled immediately, subject to the controller’s catch-up limit. With a deadline, only sufficiently recent missed occurrences are eligible. Neither mode is an unbounded or durable backlog replay. If every scheduled occurrence represents billable or compliance-sensitive work, persist those occurrences in an application queue or ledger and let Jobs claim them idempotently. CronJob history limits and events are useful operational evidence, but they are not a durable record of every intended business run.
A bounded hourly task
This example favors one active run per CronJob, permits a short controller outage window, bounds runtime and retries, and keeps failed Jobs for investigation. The image name is illustrative; replace it with an approved image in your registry.
apiVersion: batch/v1
kind: CronJob
metadata:
name: hourly-report
namespace: operations
spec:
schedule: "7 * * * *"
timeZone: "Etc/UTC"
startingDeadlineSeconds: 300
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 5
jobTemplate:
spec:
backoffLimit: 3
activeDeadlineSeconds: 900
template:
spec:
restartPolicy: Never
containers:
- name: report
image: registry.example.com/operations/report:1.4.2
args: ["run-hourly"]
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
memory: 512Mi
The history limits retain completed Job objects, not an unlimited audit trail; the documented defaults are three successful and one failed Job. Choose retention to balance diagnosis with API object growth, and export logs and business-level outcomes to durable systems. A Job’s activeDeadlineSeconds bounds total Job runtime and takes precedence over retry count. It does not undo side effects already committed before timeout.
Use explicit time and observable outcomes
CronJob timeZone is stable from Kubernetes v1.27. Specify it rather than depending on the controller manager’s local zone; do not put TZ or CRON_TZ in the cron expression. UTC avoids daylight-saving ambiguity for infrastructure schedules. If the business requirement is a local wall-clock time, use an IANA time zone and define what should happen when daylight-saving transitions create a missing or repeated local time.
Monitor schedule lateness, Job duration, failed and retried Pods, missed-start events, and the application-level count of work items completed. Starting with Kubernetes v1.32, created Jobs carry the batch.kubernetes.io/cronjob-scheduled-timestamp annotation in RFC3339 format, which helps compare intended schedule time with actual start time. Alert on stale business data, not only on Job object status: a Job can succeed while processing zero records or producing incorrect output.
Use the scheduled-timestamp annotation as correlation data, not as an idempotency guarantee: it records the controller’s intended creation time, while the application still has to atomically claim or deduplicate its work. The annotation is available only on Jobs created by CronJobs on v1.32 and later. For older clusters, pass an explicit period or work key into the task when possible, or derive the interval from durable application state. Keep both the Kubernetes Job UID and business work key in logs so retries and duplicate Job objects can be distinguished during recovery.
Before rollout, validate the rendered resource against the target cluster, then test slow execution, transient failure, process termination after an external request, controller delay, and manual invocation. Confirm that a retry does not duplicate effects, that an overlapping trigger is safe, and that operators can identify both the scheduled occurrence and its final business result.
Related:
- How the Kubernetes Scheduler Actually Places Workloads
- Kubernetes Controller Reconciliation: Idempotent Desired-State Loops
Sources: