Skip to content
SRE & DevOpsDeep Dive Published Updated 9 min readViews unavailable

AWS EventBridge Scheduler Reliability: Retries, Time Zones, and DLQs

Design scheduled AWS workloads around real timing guarantees, at-least-once delivery, bounded retries, dead-letter recovery, and observable target completion.

Amazon EventBridge Scheduler is a managed service for invoking an AWS API on a recurring rate or cron expression, or once at a specified time. It is useful for recurring maintenance, delayed work, account automation, and launching workflows without running a scheduler process in an application cluster. The key operational distinction is that a scheduled invocation is a delivery attempt, not proof that the business operation completed successfully.

Build the design around four separate contracts: when Scheduler may attempt delivery, how it retries a failed target request, what the target considers successful, and how operators find work that was accepted but later failed. If those boundaries are not explicit, an apparently simple nightly job can silently skip local-time transitions, run twice after a retry, or appear successful while its asynchronous consumer is still failing.

Select a schedule that matches the clock you mean

Scheduler supports rate-based, cron-based, and one-time schedules. A rate expression describes elapsed time; a cron expression describes calendar fields and can be evaluated in a named time zone. Choose based on the business rule, not on which expression is shorter. For example, “every 24 hours” and “at 02:30 local time every day” are not the same requirement around daylight-saving transitions.

Without a flexible time window, invocation timing is still documented with 60-second precision. A schedule set for 03:15 can be invoked any time from 03:15:00 through 03:15:59; it is not an exact-to-the-second trigger. When an invocation may happen any time within a bounded interval, configure a flexible window deliberately; Scheduler can spread invocations within that window rather than concentrating a large fleet of jobs on the same instant. Do not enable a window for a deadline-sensitive task unless the business owner accepts the maximum delay.

Time zones matter for cron and one-time schedules; Scheduler uses the IANA time-zone database to evaluate them. AWS documents the daylight-saving rules for cron schedules: if a spring-forward transition removes the configured local time, that occurrence is skipped; if a fall-back transition repeats a local time, Scheduler invokes it once rather than twice. A daily rate expression such as rate(1 days) represents a 24-hour interval, not “the same wall-clock time tomorrow”. Use cron with an explicit time zone for the latter, and define what the application should do on a skipped local-time occurrence.

aws scheduler create-schedule \
  --name nightly-reconcile \
  --group-name production \
  --schedule-expression 'cron(15 3 * * ? *)' \
  --schedule-expression-timezone 'Etc/UTC' \
  --flexible-time-window '{"Mode":"OFF"}' \
  --state ENABLED \
  --target '{"Arn":"arn:aws:lambda:REGION:ACCOUNT:function:reconcile","RoleArn":"arn:aws:iam::ACCOUNT:role/scheduler-reconcile","Input":"{\"schemaVersion\":1,\"operation\":\"reconcile\"}","RetryPolicy":{"MaximumEventAgeInSeconds":1800,"MaximumRetryAttempts":5},"DeadLetterConfig":{"Arn":"arn:aws:sqs:REGION:ACCOUNT:reconcile-schedule-dlq"}}'

Replace the ARN placeholders with resources from the same intended deployment environment. The expression above is UTC by choice; if a business schedule is local, select its IANA time-zone name explicitly and test its daylight-saving behavior. Keep this configuration in infrastructure-as-code, not as an undocumented console-only setting.

Treat delivery as at least once

EventBridge Scheduler provides at-least-once delivery to targets. A target can therefore receive more than one attempt for the same scheduled business occurrence. Make the operation idempotent where possible: derive a stable key for the intended period or object, persist whether that work has already been applied, and make replay safe. A unique execution ID is useful for tracing a delivery attempt, but it is not a substitute for an application-level idempotency key shared by retries of the same business operation.

The target’s success response also has a scope. For an SQS target, a successful SendMessage means the queue accepted the message; it does not mean a consumer completed the work. For Lambda, Scheduler invokes the Lambda API asynchronously, so later function execution errors belong to Lambda’s asynchronous processing and monitoring path. A workflow launched by Step Functions has its own execution status. Monitor the downstream service as well as Scheduler, and do not infer completed business work from a successful handoff.

Include schedule identity, scheduled time, attempt number, and a business correlation key in the target input or logs. Scheduler supports context attributes such as <aws.scheduler.schedule-arn>, <aws.scheduler.scheduled-time>, and <aws.scheduler.attempt-number>. These let responders correlate a delayed or repeated attempt with the intended occurrence. Avoid using only the current wall-clock time at the target: retries can happen later, but the original scheduled time is what explains why the work exists.

Bound retries by the usefulness of late work

A retry policy has two limits: maximum event age and maximum retry attempts. Scheduler retries with delayed attempts until it reaches either configured limit. The API accepts a maximum event age from 60 through 86,400 seconds and from 0 through 185 retry attempts. These are upper bounds, not recommendations. For an hourly cache refresh, a delayed retry may still be useful; for a time-sensitive auction close, processing an old occurrence much later could be incorrect. Set the bounds from the work’s lateness budget, target recovery characteristics, and duplicate-handling design.

Retries do not fix permanent errors such as an invalid target ARN, malformed input, missing permissions, or a target that was deleted. Validate target configuration and input before enabling a schedule, then test a representative invocation. Universal targets can call many AWS API operations, but Scheduler does not validate every Input field against the target API at schedule-creation time; some errors only surface during delivery.

When target delivery exhausts its retry policy, a configured SQS dead-letter queue (DLQ) can retain a diagnostic payload with invocation details and target response data. Give Scheduler the permission required to send to that queue, and make sure the queue policy and encryption configuration permit the intended delivery. A DLQ is an investigation and recovery boundary, not an automatic retry loop: establish an owner, retention policy, alert, triage procedure, and an explicit replay decision. Replaying a message should preserve idempotency and audit context. First fix the cause, then decide whether the original business action is still safe and timely.

Monitor both successful DLQ delivery and failure to deliver to the DLQ. If the DLQ is missing, inaccessible, or misconfigured, the fallback path can fail too. Scheduler emits CloudWatch metrics including InvocationAttemptCount, TargetErrorCount, InvocationDroppedCount, and, for schedules with a DLQ, InvocationsSentToDeadLetterCount and InvocationsFailedToBeSentToDeadLetterCount. Alert on dropped work and failed DLQ delivery, and use the DLQ payload’s error code and exhausted-retry condition to distinguish an exhausted retry budget from an expired event age.

Scope execution permissions to the target

Each schedule uses an execution role that Scheduler assumes to call its target. Grant only the API actions and resource ARNs needed for that target; a schedule that only invokes one Lambda function should not receive broad permissions over every function or service. Keep the schedule-management identity separate from the runtime execution role so a deployer can configure schedules without accidentally giving the invoked workload permission to manage them.

For the trust policy, follow AWS’s EventBridge Scheduler confused-deputy guidance and scope the aws:SourceArn condition to the intended schedule group ARN, with the account condition where appropriate. AWS specifically cautions against scoping this condition to one schedule ARN or name prefix. This is a configuration detail worth checking in infrastructure review because a trust policy that is too broad can allow a different schedule in the account to use the role. Also account for iam:PassRole in the deployment principal: schedule creation needs to pass the selected execution role, but the deployer does not need unrestricted role-passing rights.

One-time schedules need lifecycle management

A one-time schedule does not disappear just because it has invoked its target. Completed one-time schedules still count against the account’s schedule quota until deleted. Use the supported delete-after-completion action for ephemeral schedules or run a reconciler that removes completed records after preserving the audit evidence you need. For recurring schedules with an end date, automatic deletion occurs after the last planned invocation; without an end date, do not assume it will clean itself up.

Be careful when disabling and re-enabling a one-time schedule. If it is re-enabled after its original scheduled time has passed, it may invoke the target immediately. A pause button is not necessarily equivalent to cancelling a business operation. For a cancellation workflow, update the authoritative work item as cancelled and decide whether the schedule should be deleted, disabled, or safely allowed to fire a no-op.

When updating schedules through UpdateSchedule, send the full desired configuration. The API requires the schedule’s required parameters, and fields omitted from the request can be reset to null. Treat an update as replacement of the schedule definition, not a small patch, and keep the schedule expression, time zone, flexible window, state, target, retry policy, and DLQ together in reviewed declarative configuration.

Observe delivery and downstream completion separately

Create alarms for missing or failed invocations according to the workload’s expected frequency. A sparse daily schedule may not produce a useful signal from a generic service-health dashboard; compare attempts and successful application work against the specific expected schedule. For high-impact jobs, emit an application metric only after the business operation reaches its completion condition, and alert when the expected completion has not arrived by its deadline.

For diagnosis, inspect the schedule’s enabled state, expression and time zone, last intended time, target ARN, role permissions, retry age and count, DLQ metrics, and target-side logs. A Scheduler metric can confirm delivery behavior, while Lambda errors, queue age, Step Functions execution status, or application records confirm what happened after handoff. Preserve the event and deployment revision during an incident so responders can tell a bad expression from a bad target, a permissions failure, throttling, or an application defect.

Before production, exercise a controlled schedule against a test target; verify its time zone and next-fire expectation; induce a retryable target error; confirm the DLQ path; and replay one message through the normal idempotent handler. Test a one-time schedule’s deletion and the behavior of disabled schedules past their intended time. The schedule is production-ready only when both its timing contract and its recovery path are observable and owned.

Related:

Sources:

Comments