Alertmanager High Availability: Peering, Deduplication, and Delivery
Build reliable Alertmanager HA with direct Prometheus fan-out, a healthy peer mesh, stable alert identity, tested routing, and realistic delivery guarantees.
Alertmanager high availability is not achieved by placing two instances behind one load balancer. Prometheus should send alerts to every Alertmanager instance, while the Alertmanagers form a peer cluster that shares silences and notification history. This design preserves independent alert ingestion and lets the instances coordinate notification deduplication without turning one network endpoint into a single point of failure.
The delivery contract is deliberately not exactly once. Alertmanager favors at-least-once notification delivery, so a network partition or ambiguous receiver timeout can produce a duplicate notification. A correct production design makes duplicates safe and actionable, keeps alert identity stable across redundant Prometheus senders, and monitors the path from rule evaluation through the human-facing receiver.
Separate the three availability boundaries
There are three cooperating but independent systems:
- Prometheus evaluates rules. Each Prometheus server evaluates its configured rules and sends firing and resolved alert updates to Alertmanager. High-availability Prometheus replicas may evaluate equivalent rules independently.
- Alertmanager groups and routes. It matches labels against a routing tree, groups related alerts, applies silences and inhibition, and creates notification attempts for configured receivers.
- The receiver delivers to people or systems. Email, chat, paging services, and ticket systems have their own API limits, availability, retries, and acknowledgement semantics.
A green Prometheus target does not prove that rules loaded. A firing alert in Prometheus does not prove that Alertmanager received it. An alert visible in Alertmanager does not prove a route matched the intended team or that a receiver accepted and delivered the notification. Make each boundary observable and testable.
Alertmanager instances communicate through a gossip peer mesh. They replicate silences and notification-log state so that multiple instances can suppress redundant sends during normal operation. The replicated state is eventually consistent, not a consensus log. If a partition prevents peers from sharing recent notification history, each side may send. That is the expected fail-open tradeoff: duplicate an alert rather than suppress a potentially critical notification because coordination is unavailable.
Fan alerts out to every Alertmanager
Configure Prometheus with all Alertmanager targets. Do not route Prometheus through a load balancer that chooses only one instance per alert stream; the chosen target can become a bottleneck or unavailable endpoint, and the other Alertmanagers will not receive the same alert updates directly.
alerting:
alertmanagers:
- static_configs:
- targets:
- alertmanager-0.monitoring.svc:9093
- alertmanager-1.monitoring.svc:9093
- alertmanager-2.monitoring.svc:9093
This is illustrative configuration. Use the discovery mechanism and API version supported by the deployed Prometheus version and platform. Ensure each target resolves to a distinct Alertmanager process, the client can reach the configured HTTP endpoint, and network policy allows the request. A service name that load-balances to one arbitrary replica is not equivalent to listing all instances.
For a redundant Prometheus pair, labels also matter. Each sender commonly adds an external label that identifies its replica. If that replica-only label remains on the outgoing alerts, Alertmanager sees different label sets and therefore different alert identities. Use alert relabeling to remove only the replica-distinguishing label when the two servers represent the same logical monitoring system and evaluate equivalent rules:
global:
external_labels:
cluster: production-east
prometheus_replica: prometheus-0
alert_relabel_configs:
- action: labeldrop
regex: prometheus_replica
The second Prometheus replica can set the same stable cluster label and a different prometheus_replica value; both then send matching alert label sets after this narrow relabel rule. Do not remove the cluster, environment, tenant, or service identity labels that distinguish independent alerts. Do not use the same identity for unrelated clusters. Verify the resulting labels on actual alerts in Alertmanager before relying on deduplication.
This technique only makes equivalent alerts identical at the Alertmanager boundary. It does not make two rule evaluations equivalent if their expressions, labels, external labels, or alert timing differ. Keep rule files and relevant configuration aligned, compare firing states during rollout, and preserve enough source-side identity to diagnose which Prometheus replica produced an update.
Form a peer mesh deliberately
Each Alertmanager instance needs a stable, reachable address that its peers can use. Configure the cluster flags or deployment mechanism with all peer addresses, and confirm that the advertised address is routable from every member. Kubernetes deployments commonly use stable per-Pod DNS names and a headless peer-discovery service, but the exact service and chart settings depend on the operator or distribution in use.
alertmanager --cluster.listen-address=0.0.0.0:9094 \
--cluster.peer=alertmanager-0.monitoring.svc:9094 \
--cluster.peer=alertmanager-1.monitoring.svc:9094 \
--cluster.peer=alertmanager-2.monitoring.svc:9094
The peer port is separate from the HTTP API and receiver integrations. Allow the required cluster traffic between all members in both directions according to the deployed version and transport configuration. Do not expose the peer listener broadly to untrusted networks. In managed or containerized environments, inspect the effective arguments, DNS records, advertised addresses, and network policy rather than assuming a Helm value was applied.
When a member starts or rejoins, it may wait for the gossip state to settle before processing notifications. A newly healthy process can therefore take longer to resume normal notification processing than a simple TCP readiness check suggests. A partial mesh, unreachable peer address, or blocked gossip traffic can leave members with different views of silences and notification history. Alert on peer membership and failed peers, and use a canary notification during planned topology changes.
High availability is not the same as a quorum-based consensus system. Do not apply etcd-style majority assumptions to Alertmanager gossip. More peers can improve resilience to individual process or host failures, but independent network partitions can increase duplicate notifications. Choose topology and failure domains based on the receiver’s tolerance for duplicates and the consequences of notification delay.
Design routing and grouping as an operator interface
Alert identity comes from the full label set. Notification grouping is a separate configuration that chooses selected labels, such as cluster and alertname, to combine related alert instances into one message. A group should reduce noise without hiding the actionable dimensions needed to find affected services or instances. If grouping omits an important label, include that label in the notification template or keep separate groups where operators need independent pages.
The timing options have different responsibilities:
group_waitdelays the first notification for a new group, allowing nearby alerts and inhibiting alerts to arrive before the first page. Too long delays urgent incidents; too short can send incomplete pages or miss an inhibition that has not arrived yet.group_intervalcontrols when Alertmanager checks an existing group for new or resolved members. It also bounds the notification pipeline context, so a value shorter than a slow receiver’s response time can cancel sends.repeat_intervalcontrols reminders for an unchanged firing group. It is checked at group intervals, so a non-multiple is rounded up; it is also constrained by Alertmanager data retention.
Defaults are not a service-level objective. Derive timing from the incident response target and measured receiver latency. Use different child-route timing only when the urgency and receiver behavior justify it. Validate that a route’s matchers are mutually understandable, that unmatched alerts reach a deliberate fallback receiver, and that continue behavior does not accidentally send one alert to multiple paging systems.
Silences and inhibition are also operational controls, not proof that an alert is healthy. Test that expected labels match, that a broad silence cannot hide unrelated teams, that an inhibition rule has equal labels for the intended scope, and that silence changes are visible on all peers. Prefer narrowly scoped, expiring silences with a recorded owner and reason.
Persistence, reloads, and safe change
Alertmanager persists its silence and notification-log state locally and replicates that state through the peer mesh. Alerts themselves are not a durable queue: Prometheus periodically resends firing alert updates. Provide persistent storage appropriate to the deployment, preserve state during planned restarts, and confirm the peer mesh can restore shared context after replacement. A newly created empty volume can lose local history even if another peer later propagates some state.
Validate configuration before rollout with the version-matched amtool check-config or the validation path provided by the operator. Configuration files can reload at runtime through SIGHUP or the POST /-/reload management endpoint; invalid reloads are rejected and logged. After a reload, inspect active configuration and exercise a representative route. A successful parser check does not test receiver credentials, network access, templates, or actual delivery.
Use a staged change: validate syntax, apply it to one instance or a canary environment, confirm configuration load and peer health, send a controlled test alert, then roll out to the rest of the group. When changing labels or grouping, compare before-and-after alert identities and groups. Removing a label can merge alerts that were previously distinct; adding a replica label can split one logical alert into duplicates.
Test the complete notification path
Create a test alert with known labels that match a non-paging test route. Confirm that both Prometheus replicas expose the same logical alert, every Alertmanager receives it, the intended group appears, and the expected receiver records delivery. Then test resolution, repeat behavior, a silence, an inhibition, a receiver timeout, and one Alertmanager restart. A controlled receiver or staging integration is safer than sending test pages to a live on-call rotation.
For high availability, temporarily remove one instance from the peer mesh in a controlled environment and verify that Prometheus still reaches the remaining members. Test a network partition separately: duplicates are possible and should be understood, but the alert must not silently disappear. Restore connectivity and check that peers converge on the same silence and notification state. Never infer partition behavior from a process-kill test alone; network loss and process loss exercise different failure modes.
Monitor at least:
- Alertmanager readiness and HTTP/API errors, separately from process liveness.
- Cluster membership, failed-peer signals, and whether every expected member is visible.
- Alert notifications attempted, failed, retried, and sent, using metrics available in the deployed release.
- Prometheus alert delivery errors and the number of configured/reachable Alertmanager targets.
- Receiver-side API acceptance, rate limiting, delivery latency, and final acknowledgement where available.
- Firing alerts that remain unresolved and test notifications that do not reach the expected destination.
Alertmanager’s /-/healthy endpoint is a liveness-style check that always returns success while the HTTP process responds. /-/ready indicates that it is ready to serve queries. Neither endpoint proves peer convergence or receiver delivery. Pair them with cluster metrics, Prometheus send errors, and an external end-to-end synthetic test.
Production acceptance checklist
- Prometheus sends alerts directly to every Alertmanager instance, not through a single load-balanced target.
- Alertmanager peers have stable, reachable addresses and the required peer traffic is allowed only between intended members.
- HA Prometheus replicas use stable common alert identity; only the replica-specific label is removed when the rules truly represent the same alert source.
- Route grouping, fallback receivers, inhibition, silence scopes, and repeat timing are tested with known labels.
- Persistent state and behavior after restart, scale-up, and peer loss are documented and exercised.
- Configuration changes are version-validated, rolled out safely, and tested through a non-paging receiver.
- Duplicate delivery is explicitly accepted and handled; exactly-once delivery is not assumed.
- Monitoring covers rule evaluation, Prometheus-to-Alertmanager transport, peer health, and receiver delivery as separate stages.
An Alertmanager cluster is reliable when its failure behavior is understood, not merely when multiple Pods are Running. Direct sender fan-out, stable alert labels, a healthy mesh, disciplined routing, persistent state, and end-to-end tests together reduce the chance that a rule becomes either a missed page or an unhelpful duplicate.
Related:
- Prometheus Alert Rules: Evaluation, Pending State, and Notification Boundaries
- Prometheus Remote Write: Backpressure, WAL Recovery, and Delivery
Sources: