Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

Packer Machine Image Pipelines: Build, Test, and Promote Immutable Images

Build, test, distribute, and promote Packer machine images with pinned inputs, launch checks, regional validation, and an auditable rollback path.

HashiCorp Packer builds machine images from a declared base and a sequence of provisioning steps. An AMI or cloud image can then be the versioned input for a fleet, reducing the amount of configuration that must happen during every instance boot. The useful production outcome is not merely a successful packer build; it is a specific image artifact whose inputs, tests, regions, consumers, and retirement path are known.

Treat image creation as a release pipeline. Pin the builder plugin and base-image identity, validate the template, create a new artifact instead of mutating a shared image, launch a disposable instance to test the result, and only then promote the artifact for infrastructure consumers. A build can exit successfully while producing an image that cannot boot, has incomplete initialization, or fails the workload’s real readiness checks.

Keep the template declarative and dependencies explicit

Packer’s HCL template separates a source (the platform-specific build input), a build (which sources to run), provisioners (how to configure the temporary machine), and optional post-processors (what to do with the resulting artifact). Keep these concerns readable and put environment-specific values in declared variables. Avoid hiding important behavior in a large shell script that downloads unpinned content at build time; when scripts are necessary, version and test them like application code.

HCL templates can declare plugins in required_plugins. Run packer init as a normal CI preparation step so the required plugins are installed, and use a reviewed version constraint compatible with the template. Do not rely on whichever plugin happens to be preinstalled on a developer’s workstation. HashiCorp recommends vetting third-party plugins; for repeatable builds, record the Packer CLI and plugin versions used by the successful release and make upgrades an intentional change.

packer {
  required_plugins {
    amazon = {
      source  = "github.com/hashicorp/amazon"
      version = "~> 1.0"
    }
  }
}

variable "source_ami_id" {
  type        = string
  description = "Reviewed base AMI ID for this build region."
}

variable "release_id" {
  type        = string
  description = "Immutable CI release identifier."
}

source "amazon-ebs" "service" {
  region        = "us-east-1"
  source_ami    = var.source_ami_id
  instance_type = "t3.small"
  ssh_username  = "ec2-user"
  ami_name      = "service-${var.release_id}"

  tags = {
    Application = "service"
    Release     = var.release_id
    ManagedBy   = "Packer"
  }
}

build {
  sources = ["source.amazon-ebs.service"]

  provisioner "shell" {
    script = "scripts/install-service.sh"
  }
}

This is an illustrative shape, not a copy-and-run AMI template: set the SSH username and source image for the selected operating system, add the required network and instance settings, and decide how package versions are selected. Pinning a base AMI ID gives the build a stable starting point; if policy requires a moving approved base, resolve it in a reviewed update process and record the resolved ID in the build metadata. A mutable query such as “latest image” can make two builds from unchanged template text have different inputs.

Validate in stages, with clear limits

Use a staged CI flow so cheap failures are caught before cloud resources are created:

packer fmt -check .
packer init .
packer validate -var-file=ci.pkrvars.hcl .
packer build -var-file=release.pkrvars.hcl .

packer validate checks template syntax and configuration and returns a non-zero status when validation fails. It is not an integration test of the resulting image. Some validation options can evaluate data sources, which may call external services; use that only when the CI identity, cost, and side effects are understood. After the build, launch the artifact in an isolated test environment and verify boot completion, expected service enablement, network access, monitoring enrollment, and the workload-specific health endpoint. Run tests from the same instance family and subnet class used in production where those differences affect behavior.

Define the image acceptance contract before building. A practical contract might state that the instance launches from the new AMI, reaches the expected management channel, has the intended service and configuration, emits logs and metrics, passes an application smoke test, and shuts down cleanly. Validate both the first boot and a reboot. A Packer provisioner can create files and install packages during the image build, but it cannot prove cloud-init, instance metadata, production IAM, private package mirrors, or runtime network paths work after launch.

If a provisioner installs operating-system packages, keep package inputs and repositories controlled enough to explain later why two image builds differed. Clean package caches and temporary files only after the install and verification steps have succeeded. Do not bake environment-specific credentials, instance identity, private keys, or live service endpoints into a reusable base image; inject environment-specific configuration through the runtime’s managed mechanism.

Publish evidence with the artifact

A build identifier should let an operator map a running instance back to the image and source revision that created it. Record the AMI ID, source AMI ID, Packer version, plugin version, Git commit, build timestamp, test result, and relevant dependency versions. Apply useful metadata to the AMI and snapshots, and emit a manifest from the Packer manifest post-processor when a machine-readable record of build artifacts is useful. Keep the manifest tied to the CI run and do not assume that a build log alone is a durable release record.

Packer’s Amazon EBS builder can copy an AMI to other regions. Treat each regional copy as a deployment artifact that must finish and be verified before a consumer depends on it; copies can take many minutes. Check that the destination-region encryption key, launch permissions, snapshot access, and regional naming conventions are correct. A successful build in one region does not imply that an identical, usable image is available everywhere.

With HCP Packer, the registry stores metadata about images, including creation information and build provenance; it does not store the AMI or other artifact itself. Buckets organize image families, versions capture immutable build metadata, and channels are human-readable references consumers can use for stages such as testing and production. Promote a known, validated version by moving its channel only after image tests and deployment checks pass. Because a channel can point to a newer version later, production rollouts should record the resolved artifact identity for each deployment rather than relying only on the channel name as historical evidence.

Channels improve discoverability and controlled selection; they do not replace cloud permissions, retention policy, or a rollback test. Before revoking or deleting an image, determine whether launch templates, Auto Scaling groups, disaster-recovery plans, or another account still depend on it. Keep the last known-good artifact available for the full rollback window and rehearse launching it in the regions where it is needed.

Consume immutable image references

Downstream infrastructure should use the exact AMI ID or a release channel resolved during planning, not an unreviewed mutable lookup embedded in every apply. If Terraform obtains an image through an HCP Packer data source, review the selected fingerprint and regional artifact as part of the plan. A stable AMI reference makes a deployment auditable, but it does not freeze everything an instance later downloads at boot; separate the baked image contents from runtime configuration and package updates.

Roll out image changes gradually using the compute platform’s supported mechanism. Maintain compatibility between old and new images while instances overlap, and do not combine an irreversible data migration with a host replacement unless the migration is separately staged. If the canary is healthy, continue promotion; if it fails, restore the previous image reference and verify that the platform actually replaces or rolls back the affected instances. A plan that merely changes a launch template does not guarantee existing instances were replaced.

Control cost and clean up safely

Image builds create temporary compute instances, storage snapshots, regional copies, and sometimes failed intermediate artifacts. Set ownership tags, maintain an inventory of source and output artifacts, and periodically review unreferenced images and snapshots. Cleanup should be based on known consumers and retention requirements, not only an age threshold. The newest image is not automatically safe to delete from every rollback path, and an unreferenced snapshot may still be part of a retained AMI.

Keep build concurrency bounded. Parallel image builds can consume quota, network bandwidth, and package-mirror capacity; they can also race if they share an output name or mutable registry channel. Use release-scoped image names, publish metadata after a successful build, and promote a channel as a separate controlled stage. If a build fails after creating resources, use the build logs and cloud inventory to identify leftovers before retrying rather than assuming every partial artifact was removed.

Release checklist

  • Packer CLI, required plugins, source image identity, and relevant package inputs are recorded and reviewed.
  • packer fmt, packer init, and packer validate run in CI with explicit variable files and a least-scoped cloud identity.
  • The completed image is launched and tested independently of the build machine, including a reboot and workload smoke test.
  • AMI IDs, source revisions, test results, regional copies, and registry metadata are retained as release evidence.
  • Consumers use a reviewed artifact identity, and production promotion occurs only after acceptance checks.
  • A prior image remains available through the rollback window, and the recovery path has been exercised.
  • Image and snapshot cleanup accounts for actual consumers, disaster recovery, and retention obligations.

Packer turns machine-image configuration into code, but image quality depends on the evidence around the artifact. A robust pipeline records what was built, proves that it boots in the intended environment, promotes a known version, and preserves a tested route back to the prior release.

Related:

Sources:

Comments