Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

AWS Lambda Concurrency: Capacity Reservations, Cold Starts, and Throttling

Size Lambda concurrency from measured duration and demand, distinguish reserved from provisioned capacity, and diagnose throttles with the right metrics.

AWS Lambda concurrency is the number of function invocations in flight at a point in time. Each concurrent request needs execution capacity, so concurrency directly affects whether requests run, wait in an event source, spill over to a cold environment, or receive a throttle response. Production planning requires separating three concepts that are often conflated: the account’s concurrency quota, a function’s reserved concurrency, and provisioned concurrency attached to a version or alias.

Concurrency is not a requests-per-second setting. A useful starting estimate is average request rate multiplied by average execution duration in seconds. If a function receives 80 requests per second and takes 0.25 seconds on average, the steady-state concurrency estimate is about 20. That is not a safe capacity ceiling: bursts, tail latency, retries, fan-out, downstream limits, and duration variance can increase the peak. Use measured percentiles and load tests for sizing rather than treating the average formula as an SLA.

Reserved and provisioned concurrency solve different problems

Reserved concurrency allocates a portion of a regional account concurrency pool to one function and sets its maximum concurrency. This protects a critical function from being starved by other functions, while also preventing that function from exceeding a boundary that could overwhelm a database or downstream API. Reserved capacity is not pre-warmed. If a function has reserved concurrency, that allocation is unavailable to other functions even when it is idle.

Provisioned concurrency initializes execution environments in advance for a published function version or alias. Its purpose is predictable initialization latency for the allocated baseline, not an unlimited capacity guarantee. When demand exceeds the provisioned amount, Lambda may use additional reserved or unreserved capacity if available, and those invocations can incur ordinary initialization latency. If the function reaches its reserved maximum or the regional pool is exhausted, additional requests can be throttled.

When both settings are used, the provisioned value must fit within the function’s reserved capacity. Reserved concurrency caps total function usage; provisioned concurrency is a preinitialized subset of that cap, not an additional reservation on top of it. A function with reserved concurrency consumes that allocation from the regional pool even while idle; a function without reserved concurrency consumes regional capacity for its provisioned allocation. Provisioned concurrency can also incur charges even when unused. Model both the account-capacity accounting and cost before allocating high baselines across many aliases.

# Reserve an upper concurrency boundary for the function.
aws lambda put-function-concurrency \
  --function-name invoice-worker \
  --reserved-concurrent-executions 120

# Pre-initialize a baseline for the published alias used by production.
aws lambda put-provisioned-concurrency-config \
  --function-name invoice-worker \
  --qualifier live \
  --provisioned-concurrent-executions 40

# Inspect the alias-level provisioned configuration.
aws lambda get-provisioned-concurrency-config \
  --function-name invoice-worker \
  --qualifier live

The commands change live configuration; run them only through an approved deployment process after checking the account, region, function, alias, and downstream capacity. The sample values are illustrative. Provisioned concurrency is associated with a published version or alias, not $LATEST; deploy the version and update the alias deliberately before allocating it. Rollback should include returning the alias to the prior version and restoring its capacity policy.

Account for event-source behavior

An API Gateway or synchronous caller experiences a throttle directly unless its client retries. Asynchronous Lambda invocation queues events and has retry and dead-letter behavior. SQS event source mappings poll messages and scale processing according to mapping configuration and available function concurrency. Kinesis and DynamoDB Streams have shard and ordering characteristics that change how concurrency limits affect throughput. A concurrency number alone does not describe queue age, event retries, duplicate delivery, or data loss risk.

For SQS, configure event-source maximum concurrency where supported to prevent one mapping from consuming capacity needed by other workloads. Ensure the function’s reserved concurrency is not lower than the total intended mapping concurrency unless throttling is an explicit backpressure control. For stream sources, understand per-shard concurrency and ordering before enabling parallelization. In all cases, handlers should be idempotent where retries can redeliver an event, and downstream systems should have their own connection and request limits.

Reserved concurrency can be a safety brake for a function that might otherwise flood a database. It can also cause a backlog or retry storm if set below normal demand. Choose it with queue retention, visibility timeout, retry schedule, and service objectives in mind. A throttle can be safer than saturating a database, but only if the upstream event path reliably retains and replays work.

Measure concurrency and distinguish capacity signals

ConcurrentExecutions measures active invocations and is emitted at minute granularity. UnreservedConcurrentExecutions helps show demand consuming the shared pool. ClaimedAccountConcurrency reflects capacity unavailable to on-demand invocations, including allocated reservations, so it is often a better view of remaining pool than subtracting active executions from a quota. For provisioned workloads, inspect ProvisionedConcurrentExecutions, ProvisionedConcurrencyInvocations, ProvisionedConcurrencySpilloverInvocations, and ProvisionedConcurrencyUtilization with the appropriate statistic.

Interpret metrics as a set. A rising ProvisionedConcurrencySpilloverInvocations count means some invocations ran outside the preinitialized allocation, but it does not necessarily mean they were throttled. High utilization with low spillover may be expected, while low utilization over a long period can indicate excess provisioned capacity. A function’s Throttles metric, errors, duration percentiles, event-source age, and downstream latency provide the broader picture.

aws cloudwatch get-metric-statistics \
  --namespace AWS/Lambda \
  --metric-name ConcurrentExecutions \
  --dimensions Name=FunctionName,Value=invoice-worker \
  --start-time 2026-10-04T00:00:00Z \
  --end-time 2026-10-04T01:00:00Z \
  --period 60 \
  --statistics Maximum

This example requests a one-hour maximum series; confirm the time range and region before interpreting it. CloudWatch metric dimensions and availability differ by metric and alias/version. For operational dashboards, use the metrics dimensions documented for the selected function and alarm on a service-level objective such as sustained throttles, growing queue age, or latency rather than an arbitrary concurrency number alone.

Diagnose throttling in layers

First identify whether the throttle is due to the function’s reserved limit, regional account concurrency, a provisioned-concurrency boundary, or an event-source mapping setting. Compare the function’s active concurrency with its reserved maximum and review regional ClaimedAccountConcurrency and quota. Then inspect the calling model: synchronous clients may need controlled backoff, whereas queue-based systems need age and retry analysis. Do not raise the regional quota until confirming that downstream dependencies can accept the added parallelism.

If provisioned spillover rises but no throttles occur, decide whether the added cold-start latency violates the workload objective. Provisioned capacity is tied to a version or alias, so verify that production traffic actually uses the alias with the allocation. Check whether provisioned allocation completed successfully before shifting traffic. New versions do not automatically inherit the desired concurrency setting unless your deployment process configures it.

If a function reaches reserved concurrency, review duration, request rate, timeout, memory, and dependencies. Raising concurrency can reduce backlog but may increase concurrent database connections, API rate limits, or costs. Consider reducing execution duration, batching work, applying a token bucket to a downstream integration, or using a queue to absorb bursts. Tune memory only with measurement; it can affect CPU allocation and duration, but it is not a direct concurrency cap.

Deploy concurrency changes safely

Record the current function alias, reserved setting, provisioned setting, account quota, and expected peak before a change. Apply one capacity variable at a time and observe metrics through a representative traffic window. For a launch, allocate a small provisioned baseline for the exact production alias, verify initialization, then shift traffic. For scale-down, reduce provisioned capacity before lowering the reserved cap only after confirming actual peak concurrency and event backlog.

Do not use concurrency controls as a substitute for idempotency, queue visibility design, database pooling, or request rate limiting. Reserved concurrency limits parallel function execution but cannot undo side effects after a timeout. Provisioned concurrency reduces initialization work for ready environments but does not guarantee every invocation avoids cold-start paths, and it adds cost. Keep retry policies bounded and include overload behavior in integration tests.

Effective Lambda capacity management starts with measured demand and ends with verified service behavior. Match concurrency allocation to the invocation model, account for version and alias boundaries, preserve regional headroom, and watch both spillover and downstream saturation. Every concurrency limit should have an owner, a reason, and a runbook describing which metric tells the responder to raise, lower, or leave it unchanged.

Related:

Sources:

Comments