Ansible Playbooks in Production: Idempotency, Check Mode, and Rolling Changes
Build safer Ansible operations with idempotent modules, honest check-mode limits, handler-aware recovery, bounded host batches, and health-gated rollouts.
Ansible is safest in production when a playbook describes the desired state of a host and makes the impact of each change visible. A successful run is not proof of idempotency, a --check run is not a perfect dry run, and serial is not an availability guarantee. Each is a control with specific semantics and failure modes.
This guide focuses on three operational contracts: repeated runs should converge without unnecessary work, validation should reveal what it can and clearly expose what it cannot simulate, and a multi-host rollout should stop at a bounded and observable failure point. Those contracts depend as much on task design, handler timing, host selection, and health checks as on YAML syntax.
Express desired state with modules
Most Ansible modules compare the current system with the requested state. A template task, for example, manages the contents and attributes of a file and reports a change when its rendered output differs. A service task with state: started starts a stopped service but does not issue a restart on every playbook run. These state-oriented operations can make a second identical run converge to ok instead of repeating work.
By contrast, command and shell execute operations whose effects Ansible cannot infer generically. A command that creates a user, restarts a service, or appends to a file may report changed every time, or may perform an irreversible action more than once. Prefer a purpose-built module that expresses the end state: user, package, file, template, lineinfile, service, or a provider-specific collection module.
When a command is unavoidable, explicitly define its guard and result semantics. creates and removes can make some command tasks safe to skip based on a file’s presence, but that is a coarse condition, not a general model of the external system. Use changed_when only when you can reliably determine whether the operation actually changed the host; use failed_when when the command’s return code alone does not represent success. These expressions affect Ansible’s reporting and handler notifications. They do not make a non-idempotent command idempotent.
For example, a read-only version probe should not trigger a handler:
- name: Read the installed service version
ansible.builtin.command: /usr/local/bin/orders-api --version
register: orders_api_version_result
changed_when: false
That task is only safe to classify as unchanged because it is known not to mutate state. Do not copy changed_when: false onto an installer or migration command just to get a clean recap. A false “changed” result can suppress a required handler, make a deployment dashboard misleading, and hide that the task runs every time.
Treat check mode as a partial simulation
Check mode (--check or -C) asks modules that support it to predict changes without applying them. Support varies by module. Modules with no check-mode support may skip the operation, and a task that depends on a registered result from an earlier simulated task may not have realistic input. Check mode is valuable evidence, but it cannot prove that every real task will succeed or that a non-supporting module has no side effects.
Diff mode (--diff) can show before-and-after content for supported modules, often for templates and files. It can also expose credentials, private keys, tokens, connection strings, and other sensitive values embedded in managed files. Suppress diffs on sensitive tasks with diff: false, and control where the full output is stored and who can read it. A safe command line for reviewing a single candidate host might be:
ansible-playbook deploy.yml --syntax-check
ansible-playbook deploy.yml --list-hosts
ansible-playbook deploy.yml --check --diff --limit app-01.example.net
These options answer different questions. Syntax check parses the playbook; list-hosts shows the matched target set; check and diff ask supporting tasks for predictions. Run the validations separately so each output has one clear purpose. Review the actual inventory, variables, tags, limit, privilege escalation, and connection settings that the deployment run will use. A typo in --limit or an unexpectedly broad group can matter more than a valid playbook.
Do not make a supposedly harmless task force normal execution during check mode without documenting it. check_mode: false explicitly allows a task to make real changes even when the operator uses --check. Reserve that override for a narrowly scoped operation whose side effects are understood and whose role in validation is necessary. Avoid using a check-mode run as approval to apply when critical deployment steps were skipped or their registered conditions could not be evaluated.
Use handlers as controlled side effects
Handlers run when notified by a task that Ansible reports as changed, and normally run at the end of the play for each host. This is useful for avoiding unnecessary service reloads when a template is already correct. It also creates an important failure case: if a configuration task changes a file, notifies a handler, and a later task fails on that host, the handler normally will not run there. The file may be updated while the service keeps using its old in-memory configuration.
For a change that must be activated and checked before proceeding, flush handlers at an intentional point:
- name: Install the service configuration
ansible.builtin.template:
src: orders-api.conf.j2
dest: /etc/orders-api/orders-api.conf
owner: root
group: orders-api
mode: "0640"
notify: Reload orders API
- name: Apply notified changes before readiness verification
ansible.builtin.meta: flush_handlers
- name: Verify the local readiness endpoint
ansible.builtin.uri:
url: http://127.0.0.1:8080/ready
method: GET
status_code: [200]
return_content: false
This example assumes the target runs the service and exposes that local readiness route. A real rollout may need an application-specific test, traffic drain, or external synthetic check. A successful HTTP response verifies only the behavior that endpoint actually measures. It does not prove that dependent services, database migrations, user traffic, or all application paths are healthy.
meta: flush_handlers runs pending handlers at that point, so the readiness task can observe the new service state. Use it sparingly: it changes the normal end-of-play timing for every handler notified on that host. force_handlers is another explicit option that runs notified handlers even after a later task failure, where possible, but forcing a reload is not automatically safer. Decide whether a host with partially applied configuration should restart, roll back, or stop for intervention. Model that recovery path instead of relying on a global flag without considering the service.
Keep handler names unique and make each action narrow. A restarted service action always restarts, while a reloaded action requests a reload and can start a stopped service. If a module’s semantics do not match the service manager or desired availability behavior, use the appropriate platform-specific module and test it on a disposable target first.
Bound risk with serial batches and failure policy
By default, Ansible’s linear strategy runs a task across the active host set before starting the next task. The serial keyword divides the play into batches; Ansible completes the play for the current batch before moving to the next one. This gives operators a place to observe health between groups of hosts, but it does not make an unsafe update safe by itself.
An initial batch of one can expose missing assumptions on a canary host:
---
- name: Apply the reviewed service configuration
hosts: app_servers
become: true
strategy: linear
serial: 1
any_errors_fatal: true
tasks:
- name: Install the service configuration
ansible.builtin.template:
src: orders-api.conf.j2
dest: /etc/orders-api/orders-api.conf
owner: root
group: orders-api
mode: "0640"
notify: Reload orders API
- name: Activate pending service changes
ansible.builtin.meta: flush_handlers
- name: Confirm readiness after reload
ansible.builtin.uri:
url: http://127.0.0.1:8080/ready
status_code: [200]
return_content: false
handlers:
- name: Reload orders API
ansible.builtin.service:
name: orders-api
state: reloaded
The host group, template, and health endpoint are examples that must be replaced with tested site-specific values. The play does not remove a host from a load balancer or restore it after a failure. If traffic draining is required, implement it explicitly with the load balancer’s supported API or module, wait until in-flight work is gone, and add the host back only after the new version is healthy. A one-host batch can still cause an outage if it is the only healthy instance or if capacity is already below demand.
Choose serial against the service’s redundancy, traffic, failure domain, and recovery time. A percentage is relative to the selected host set; a small percentage still executes on at least one host. For a progressive rollout, Ansible also supports a list of serial batch sizes. Do not assume a task using run_once runs exactly once across all batches: in combination with serial, it runs once per batch. Put database-wide or singleton operations in a separate play or explicitly guard them with the intended host condition.
Failure thresholds are evaluated in the context of batches. max_fail_percentage applies to each batch and the threshold must be exceeded, not merely reached. For a strict stop-on-first-fatal-task behavior, any_errors_fatal lets the current batch finish the fatal task and then stops later tasks and plays. Choose the policy that matches the operation, and make health-check failures fatal when continuing would expose more hosts to a bad release. Avoid ignore_errors or permissive failed_when logic on rollout gates unless there is a documented recovery reason.
Make rollout inputs reproducible
Ansible’s control-flow is only as deterministic as the inventory and variable inputs. Pin the inventory source and deployment revision, record the selected hosts, and make the release version an explicit input. Use --list-hosts and --list-tasks in review or preflight steps, and inspect the effective variables that select the package, template, environment, and service. Avoid relying on an operator’s current working directory, shell aliases, or unreviewed extra-vars for production identity.
Use a consistent execution environment for ansible-core and collections. A playbook that passes on one controller can behave differently if a different module version or inventory plugin is loaded. Record the Ansible version, collection requirements, configuration file, Python dependencies, and inventory plugin versions alongside the automation. Test against representative operating systems, service managers, and privilege policies instead of assuming one successful Linux distribution run proves portability.
Separate host convergence from one-time orchestration. A database migration, schema change, cache flush, or load balancer pool operation should not accidentally run once for every application host. run_once is scoped to each serial batch, so isolate singleton operations or explicitly design how they run across batches. Tasks delegated to localhost execute from the controller context; review their variables, credentials, and concurrency independently of the managed host task.
Prove convergence with repeated runs
For each playbook or role, test at least these cases in a representative non-production environment:
- A host in the original state converges to the requested state.
- A second identical run reports no unnecessary changes and does not restart services without a new notification.
- A check-mode run makes no unintended changes, and the team understands which tasks are skipped or only partially simulated.
- A diff-mode run contains no exposed secrets in logs, CI artifacts, or shared output.
- A deliberate validation failure stops the rollout according to the batch policy and leaves a known recovery path.
- The post-change health check detects a failed service or wrong configuration before another batch begins.
- The selected inventory and
--limitinclude exactly the intended hosts.
When commands must be used, test their creates, removes, return-code, and changed-state logic on both the first and repeated run. A task that always reports changed can trigger handlers forever; a task that never reports changed can conceal a required restart. A no-change second run is useful evidence of convergence, but it does not prove correctness if the playbook’s desired state is incomplete or the health check is too weak.
Production review checklist
- Prefer state-oriented modules and make every imperative task’s repeat behavior explicit.
- Verify changed and failed semantics; do not use reporting overrides to hide mutations.
- Treat check mode and diff mode as partial evidence and protect sensitive diff output.
- Confirm the exact inventory, variables, privileges, host limit, Ansible version, and collection versions.
- Use handlers to avoid unnecessary reloads, but understand their timing and failure behavior.
- Select
serial, failure thresholds, and fatal-error policy based on service capacity and blast radius. - Separate singleton operations from repeated per-host tasks, especially with
serialandrun_once. - Drain and restore traffic explicitly when the service requires it; a host batch is not an availability strategy by itself.
- Test readiness after activation and stop before expanding a failed rollout.
- Verify a repeated run converges, then test recovery and partial-failure cases in a safe environment.
Production-grade Ansible is not simply a playbook that exits zero. It is automation whose state model, validation limits, task results, batch boundaries, and failure recovery are understood by the people approving each run. Make those properties observable, then expand a deployment only when the canary and its health checks support the next step.
Related:
- How to Set Up a CI/CD Pipeline with GitHub Actions
- Kubernetes Deployment Rollouts: Progress, Availability, and Revision History
Sources:
- Ansible: Playbooks, desired state, and idempotency
- Ansible: Check mode and diff mode
- Ansible: Execution strategies, forks, and serial batches
- Ansible: Error handling, handlers, changed_when, and failure thresholds
- Ansible: Template module
- Ansible: Service module
- Ansible: URI module
- Ansible: Command module check-mode behavior
- Ansible: ansible-playbook command-line options