Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

AWS CodeDeploy EC2 Blue/Green: Traffic Cutover and Rollback

Operate CodeDeploy blue/green releases for EC2 with precise AppSpec hooks, replacement capacity, load-balancer cutover, alarm-based rollback, and safe cleanup.

AWS CodeDeploy blue/green deployments on the EC2/On-Premises compute platform are an EC2 replacement workflow: CodeDeploy prepares a replacement fleet, installs an application revision, then shifts load-balancer traffic from the original instances. The platform name is broader than this deployment type: CodeDeploy supports on-premises instances for supported deployment workflows, but AWS requires Amazon EC2 instances for EC2/On-Premises blue/green deployments. Use this model when a deployment needs a separately provisioned green fleet and an explicit traffic transition; do not confuse it with ECS or Lambda traffic-shifting configurations, which use different deployment mechanics.

The important operational decision is not merely “blue or green.” It is when the replacement instances are considered valid, how long the original fleet remains available after cutover, and what signal can stop a bad release. CodeDeploy can orchestrate those lifecycle steps, but it cannot make an unsafe schema migration reversible or prove that a host-level success response means users are receiving correct responses.

Understand the replacement sequence

A deployment group identifies the original environment through EC2 tags, Auto Scaling groups, or both. For an automatically copied Auto Scaling group, CodeDeploy provisions a replacement group based on the selected source group. It installs the requested revision on replacement instances and can pause for a configured wait period. During that period, the team can perform system verification before traffic moves. When traffic is rerouted, replacement instances are registered with the selected load balancer or target groups and the original instances are deregistered. The old fleet can then be terminated after a chosen delay or kept running.

That pause is a decision point, not an automatic test listener. If traffic rerouting is configured as manual and nobody reroutes before the wait period expires, the deployment is stopped. Plan who or what performs the verification and the transition, how its decision is recorded, and which alarm can block promotion. If you choose immediate rerouting, validate the replacement environment through deployment hooks and load-balancer health checks before relying on production traffic to reveal defects.

The original fleet’s cleanup delay is also an availability and cost decision. Keeping the instances briefly can provide rollback capacity and allow comparison; terminating them quickly reduces duplicate compute but removes a simple traffic-return option. Set an explicit delay that reflects the time needed to observe real requests and make a rollback decision, and ensure the Auto Scaling group and termination settings preserve the behavior you expect.

Make the AppSpec lifecycle intentional

For EC2 deployments, the AppSpec file is a YAML file named appspec.yml at the root of the revision bundle. It maps files into instance paths and associates scripts with lifecycle events. The CodeDeploy agent runs listed scripts sequentially within an event, and a non-zero exit marks the script as failed. A minimal example might look like this:

version: 0.0
os: linux
files:
  - source: app/
    destination: /opt/orders
hooks:
  AfterInstall:
    - location: scripts/configure.sh
      timeout: 300
      runas: root
  ApplicationStart:
    - location: scripts/start.sh
      timeout: 120
      runas: root
  ValidateService:
    - location: scripts/validate.sh
      timeout: 90
      runas: root

Choose events by responsibility. AfterInstall is a place for configuration or permission work after files are copied. ApplicationStart starts the installed service. ValidateService is the last scripted lifecycle event and should test a bounded, meaningful readiness condition. It should not only check that a PID exists if the real requirement is a successful local health endpoint or a connection to a required dependency.

ApplicationStop has a non-obvious first-deployment behavior: it uses the AppSpec and scripts from the last successfully deployed revision, because the new revision has not yet been downloaded. It does not run on the first deployment to an instance. BeforeBlockTraffic and AfterBlockTraffic also use scripts from the previous successful revision; the other script hooks use the current deployment’s AppSpec. Therefore, do not put essential initial setup only in ApplicationStop, and keep the prior revision’s stop and traffic-blocking hooks compatible with the deployment process. Test both first installation and upgrades, and specifically exercise the three prior-revision hooks with the exact version that will be installed before cutover.

Treat hooks as bounded, observable programs. Set explicit timeouts, make them safe to retry after partial progress, write useful output to logs, and fail closed when a required check is inconclusive. AWS documents a maximum of 3,600 seconds for script execution in an individual lifecycle event; the total timeouts for scripts in that event must not exceed that limit. Avoid long database migrations or unbounded polling inside a host hook. Those operations need their own concurrency, observability, and rollback plan.

Size and identify the replacement fleet

Blue/green capacity means old and new instances coexist during at least part of a deployment. Budget for that overlap in the target region, account quotas, subnet address space, EC2 capacity, Auto Scaling limits, target group registration, and any per-instance licensing constraints. If green capacity cannot be provisioned, deployment cannot reach its verification stage. Validate the Auto Scaling group launch template, AMI, instance profile, CodeDeploy agent, application dependencies, and security-group paths before release day.

Instance selection must be deterministic. Use a clear tag scheme or Auto Scaling group membership, and check what the deployment group actually resolves to before deployment. Stale instances that remain in scope can cause unexpected failures; instances accidentally excluded can leave part of a fleet on an old revision. The deployment configuration’s minimum healthy host setting applies to the replacement environment during blue/green, not to the original fleet. It also does not replace the end-to-end health check: a replacement instance can meet a host-count threshold while the application is functionally wrong.

Attach load balancers and target groups to the Auto Scaling group before starting the CodeDeploy deployment. AWS warns that attaching them after deployment creation can unexpectedly deregister instances. Check target group health-check path, port, grace periods, draining behavior, and network rules together. The health endpoint should represent readiness to serve, and the load balancer must have enough time to mark replacement instances healthy before traffic cutover.

Gate promotion with alarms and automatic rollback

Configure CloudWatch alarms that measure application outcomes during the deployment: elevated request errors, latency, failed synthetic checks, or another service-level indicator. CodeDeploy can stop a deployment when an associated alarm activates. An alarm is only useful if it has a meaningful threshold, correct dimensions, enough data points to avoid reacting to one transient sample, and permission for CodeDeploy to query its state. If alarm status is unavailable, decide explicitly whether deployment should continue or fail rather than accepting a default without review.

Automatic rollback can be configured for deployment failure and for configured alarm thresholds. A rollback deploys the last known good revision as a new deployment with a new deployment ID; it is not a rewind of the failed deployment’s side effects. Files written outside CodeDeploy’s managed revision paths, changes to databases, messages already processed, and external API calls are not undone by restoring an older application bundle. Design release artifacts to be immutable and keep stateful changes backward-compatible across the observation and rollback window.

Test rollback before relying on it. A useful rehearsal deploys a deliberately failing revision to a non-production fleet, confirms the signal stops promotion, confirms the prior revision is installed as a separate deployment, and checks load-balancer registration and user-visible behavior. Also test the no-alarm path and a broken alarm configuration. A green fleet that passes host installation but fails its application validation should never receive traffic merely because rollback is enabled.

Diagnose the lifecycle event, not just the final status

When a deployment fails, identify the first failed lifecycle event and its instance. Check deployment events, the CodeDeploy agent log, the hook’s exit code and output, instance reachability, artifact download permissions, and whether the targeted instance is still in scope. Preserve the application revision and deployment ID before retrying; a later successful deployment can obscure the partial state left on an earlier host.

For a stopped or stuck deployment, distinguish a hook that is still running from a replacement instance that never became healthy, a manual reroute that exceeded its wait period, a load-balancer deregistration failure, or a capacity/agent problem. A deployment can fail after some lifecycle events have already run, so cleanup scripts must tolerate partial completion. Do not assume CodeDeploy automatically removes every external side effect created by a hook.

Production checklist

  • Replacement capacity, subnet addresses, instance profile, agent, and artifact access have been exercised before the release.
  • The deployment group’s EC2 selection, load balancer attachments, target health checks, and minimum healthy host policy are reviewed.
  • AppSpec is at the bundle root; scripts use explicit timeouts, safe retries, useful logs, and event responsibilities that match the lifecycle.
  • First deployment and upgrade paths are both tested, including the previous revision’s ApplicationStop behavior.
  • The wait period has a named owner or an automated verification step, and the promotion decision is auditable.
  • CloudWatch alarms observe the application and are configured to stop or roll back a bad deployment.
  • The previous release artifact is immutable and database/API changes remain compatible with that artifact during rollback.
  • The original fleet retention delay is intentional, and the runbook states when to terminate it or return traffic to it.
  • Rollback, partial-hook failure, Auto Scaling replacement, and load-balancer health failure have been rehearsed in a non-production group.

CodeDeploy blue/green can make host replacement and traffic cutover repeatable, but production safety comes from matching each lifecycle signal to a real system property. Provision enough green capacity, verify it before promotion, watch the application’s behavior after cutover, and keep a tested recovery path that accounts for state outside the instance revision.

Related:

Sources:

Comments