EC2 Auto Scaling Instance Refresh: Safe AMI and Launch Template Rollouts
Roll out Auto Scaling group changes with explicit capacity bounds, warmup, checkpoints, skip-matching behavior, alarms, and rollback prerequisites.
Amazon EC2 Auto Scaling instance refresh replaces instances in an Auto Scaling group after a launch-template or mixed-instances configuration change. It is useful for rolling out a new AMI, instance configuration, or supported group desired configuration without manually terminating instances one by one. The rollout behavior is governed by healthy-capacity percentages, health checks, warmup, and optional checkpoints. These settings determine how quickly the group moves and how much capacity can be unavailable or temporarily overprovisioned.
An instance refresh is not an application-level deployment controller. It can observe Auto Scaling health checks and wait for warmup, but a successful EC2 status check does not prove that the service is processing correct requests. Pair the refresh with load-balancer health, application metrics, alarms, and a rollback plan that restores the previously tested launch configuration.
Pin the desired configuration before starting
Update the launch template to a specific numbered version and start the refresh with that desired configuration. Avoid $Latest and $Default if rollback is required: AWS documents that refresh rollback cannot target a previous launch-template version through those moving aliases. A Systems Manager Parameter Store AMI alias also makes rollback unavailable. An explicit version and AMI identifier turn the rollout into a reviewable transition from one known configuration to another.
During a concurrent scale-out event, Auto Scaling launches new instances with the refresh’s desired configuration rather than the group’s prior settings. After success, Auto Scaling updates the group settings to match that configuration. Coordinate other automation that edits the same launch template or mixed-instances policy, since changes derived from the active desired configuration can be rejected while the refresh is in progress.
{
"AutoScalingGroupName": "orders-prod",
"DesiredConfiguration": {
"LaunchTemplate": {
"LaunchTemplateName": "orders",
"Version": "42"
}
},
"Preferences": {
"MinHealthyPercentage": 90,
"MaxHealthyPercentage": 120,
"InstanceWarmup": 300,
"SkipMatching": true,
"CheckpointPercentages": [10, 50, 100],
"CheckpointDelay": 900,
"BakeTime": 1800,
"AutoRollback": true,
"AlarmSpecification": {
"Alarms": ["orders-prod-5xx-rate"]
}
}
}
This start-instance-refresh input is illustrative. Confirm every percentage, duration, launch template version, and alarm against the target group before use. The 300-second warmup, for example, must reflect measured startup and readiness time; it is not a universal recommendation. The sample also assumes a launch-template configuration compatible with skip matching. Auto Scaling groups that use attribute-based instance selection have restrictions on DesiredConfiguration and SkipMatching; check the current service limitations before applying those options.
Choose the capacity envelope deliberately
MinHealthyPercentage defines the minimum share of desired capacity that must remain healthy and ready for the refresh to continue. MaxHealthyPercentage caps how far capacity may rise during replacements. The difference between the two cannot exceed 100 percentage points. A minimum of 100 percent makes Auto Scaling launch a replacement before terminating the old instance, preserving desired capacity while replacements proceed, but requires enough EC2 quota, subnet addresses, and capacity to run temporary extra instances. A lower minimum can reduce overprovisioning at the cost of running below desired capacity during the replacement.
Do not treat these percentages as an availability guarantee. Health is determined by the health checks configured on the group, which may include EC2 status checks and load-balancer checks. A shallow probe can mark an instance healthy before the application is ready. Conversely, a broken or overly strict health check can make a good image appear unhealthy and stall the refresh. Model the minimum in instance counts for the current group size and test rounding behavior for small or weighted groups before choosing the policy.
Warmup is the interval after a new instance becomes InService during which Auto Scaling waits before advancing to the next replacement. It should cover measured boot, configuration, application initialization, and the point at which the configured health check reflects useful capacity. If the group has a correct default instance warmup, the refresh can use it unless the rollout needs an override. Avoid using an arbitrarily long value: it slows replacement and delays useful scaling metrics. An arbitrarily short value can let the next batch start before the new instances can absorb production traffic.
Use checkpoints as verification gates
Checkpoints divide a refresh into percentage-based phases. At each checkpoint, Auto Scaling pauses for the configured delay, allowing operators or automation to inspect new instances and service metrics before replacement continues. An EventBridge rule can route checkpoint events to an operations workflow or notification channel. Choose percentages that create meaningful sample sizes: in a tiny group, a 10-percent checkpoint may be crossed by replacing one instance and can be skipped together with another low threshold.
Always make the final checkpoint 100 percent when the intended outcome is a complete replacement. A refresh whose last checkpoint is below 100 percent stops at that phase instead of automatically continuing to full completion. Also distinguish cancel from rollback: cancelling or failing before the final checkpoint stops future replacements, but instances already replaced are not automatically returned to the old configuration. If the new cohort is bad, use rollback while the refresh is still eligible or start a new corrective refresh; do not assume that cancel reverses changes.
Bake time adds a final observation period after replacements, before the refresh is considered complete. Use it to detect delayed errors, saturation, or user-visible regressions that appear after a host begins serving real traffic. A bake time helps define the point at which the refresh is considered finished, but it does not run a bespoke test suite; alarms and application-level checks still need to measure the properties that matter.
Understand skip matching and console defaults
Skip matching avoids replacing instances that already match the desired configuration. Its behavior depends on whether a desired configuration is supplied and which launch-template or mixed-instance properties Auto Scaling can compare. It does not detect application-code changes that are fetched by a user-data script if the launch-template version did not change. If user data pulls a new application revision, disable skip matching or version the input so the refresh can distinguish old from new instances.
Defaults are not identical between the console and CLI/SDK. For example, current AWS documentation lists skip matching as enabled by default in the console and disabled in the CLI/SDK; scale-in-protected and standby-instance handling also differs. Minimum healthy percentage can derive from an instance maintenance policy or default to 90 percent when no policy exists. Do not rely on a remembered console default in automation. Specify the options whose behavior affects availability, and record the exact start request with the deployment evidence.
Before starting, inventory instances in Standby or protected from scale-in. Depending on the selected handling option, a refresh can wait for an operator, ignore them, or replace/terminate them. A wait can expire and fail if the condition is not resolved. Select behavior intentionally so a quiet maintenance instance is not accidentally treated as an unexplained deployment stall or terminated contrary to its operational role.
Make rollback a real, tested path
Auto rollback can restore the configuration saved on the Auto Scaling group before the refresh if the refresh fails or an associated CloudWatch alarm enters ALARM. Rollback is available only when the refresh starts with a desired configuration. The prior group configuration must also be stable; otherwise, the rollback workflow can fail and leave the group unable to launch instances. The prior launch-template version must be a specific numbered version, not $Latest, $Default, or a Systems Manager AMI alias.
CloudWatch alarms are optional, but choose signals that correspond to service health, such as load-balancer 5xx responses, failed requests, or a synthetic transaction. If an alarm is attached without auto rollback, an alarm breach fails the refresh without reversing the rollout. Ensure the alarm is in a usable state before starting; an alarm already in ALARM or INSUFFICIENT_DATA can prevent the refresh from starting. Decide how missing metric data should be interpreted, especially for low-volume services.
The rollback window is limited to an in-progress refresh. Once a refresh completes, rollback is no longer available as part of that operation; apply the known-good configuration in a new refresh. Rehearse both automatic and manual rollback before relying on either, and keep the previous AMI and numbered launch-template revision available. Rollback replaces already-updated instances with instances on the saved configuration; it does not undo application-level data changes or side effects.
Investigate stuck and partial refreshes
For a stalled refresh, read the current status and status reason before taking action. Check EC2 launch failures, subnet IP exhaustion, instance quotas, load-balancer health, health-check grace period, warmup, mixed-instance weights, scale-in protection, standby instances, and checkpoint delay. A refresh can continue retrying for a period when new instances fail health checks or protected/standby instances block replacement, then fail if it cannot recover.
Checkpoints and warmup affect what the displayed completion percentage means. After a checkpoint, progress may not update until replacement instances finish warming. Group scaling during a phased refresh can also cause a checkpoint to be reached again because percentages are evaluated against group size. This is why alarms and completion checks should use the refresh status plus application telemetry rather than treating one console percentage as the only signal.
If the operation fails after some replacements, identify which instances run the new launch-template version and which remain on the old one. Preserve the instance-refresh ID, group activity history, relevant CloudWatch alarms, and the exact desired configuration. A retry can target earlier instances first, but it does not simply resume from the visual checkpoint where the prior attempt stopped. Decide whether to roll back, continue toward the same desired configuration, or create a new version based on evidence from the partial fleet.
Operational checklist
- The launch template version and AMI are immutable, numbered, and retained for rollback.
- Minimum and maximum healthy percentages are translated into instance counts and checked against quota, subnet, and temporary capacity.
- Health checks measure application readiness, and warmup reflects measured initialization time.
- Checkpoint percentages include 100 percent; operators or automation own verification and promotion at each pause.
- Skip matching is appropriate for the deployment input; user-data-pulled changes are versioned or skip matching is disabled.
- Standby and scale-in-protected instances have an explicit handling policy.
- CloudWatch alarms are healthy before the refresh, measure meaningful service behavior, and are paired with auto rollback when rollback is intended.
- Rollback prerequisites are met, and both rollback and post-completion corrective-refresh procedures are documented.
- The release records the refresh ID, desired configuration, per-instance version outcome, alarm state, and final service verification.
An instance refresh makes a fleet change observable and paced, but its defaults are not a deployment policy. Pin the desired configuration, set the healthy-capacity and verification gates, and keep a tested path to the prior version. That turns host replacement from a batch operation into a controlled release with explicit operational evidence.
Related:
- Packer Machine Image Pipelines: Build, Test, and Promote Immutable Images
- AWS CodeDeploy EC2 Blue/Green: Traffic Cutover and Rollback
Sources: