Argo Rollouts Analysis Gates: Canary Metrics, Promotion, and Abort
Use Argo Rollouts analysis as an explicit release gate with representative Prometheus queries, bounded samples, traffic control, and rollback expectations.
Argo Rollouts provides a Rollout resource that can act as a drop-in replacement for a Kubernetes Deployment when a workload needs progressive delivery steps, traffic shaping, or metric-driven promotion. A canary is not simply “deploy fewer Pods first.” It is a controlled experiment in which the new ReplicaSet receives a defined portion of production traffic, health evidence is collected, and the release either advances, pauses for a human decision, or aborts back to the stable version.
The analysis system is only as reliable as its metrics and query. A canary query that returns no data, mixes stable and canary traffic, includes synthetic health checks but excludes real user failures, or has a denominator of zero can make a dangerous decision. Configure explicit success, failure, inconclusive, and timeout behavior; verify the query against an independently observed release; and make the rollback mechanism and traffic provider part of the design review.
Separate rollout state from traffic routing
The Rollout resource manages ReplicaSets and progression state. Traffic routing is a separate integration with an ingress controller or service mesh when precise percentages are required. Without an integrated router, the controller can approximate weights by scaling ReplicaSets, but replica ratios are coarse at small replica counts. A requested ten-percent canary with four replicas cannot be represented as accurately as a router that can distribute requests by weight.
An Argo Rollouts setWeight step expresses the desired canary traffic proportion in a supported traffic-routing integration. It should not be read as a universal guarantee of exact per-request distribution: session affinity, retries, connection reuse, routing implementation, and low traffic volume can skew observed proportions. Measure actual canary request volume and ensure the telemetry label that identifies the candidate version is attached consistently.
Stable and canary Services are selectors managed by the rollout strategy. Avoid an independent controller or operator continuously rewriting those selectors. If the traffic provider owns another routing resource, define clear ownership and inspect rendered resources after installation. Promotion and abort are state transitions that affect selectors, ReplicaSet scale, analysis runs, and possibly traffic weights; a healthy Pod alone does not prove the user path is healthy.
Design an analysis template around a measurable hypothesis
An AnalysisTemplate defines metric providers and success logic; an AnalysisRun records the evaluation during a rollout. A useful canary metric asks a specific question, such as whether the new revision’s server-error ratio remains below a limit during a representative traffic interval. Include enough samples to cover ordinary variance, and avoid a metric window so short that scrape delay or low-volume noise dominates. Include a minimum request-volume check separately when a ratio can look perfect with only one request.
The query below assumes the Prometheus series have a revision label whose value is the candidate ReplicaSet’s pod-template hash. Argo Rollouts can pass the latest hash to an AnalysisTemplate argument; when Prometheus discovers Pods, its Kubernetes service-discovery labels can be relabeled into a stable metric label such as revision. Verify that the label is present on both numerator and denominator series before using a gate. Without that instrumentation or relabeling, the query does not isolate canary traffic.
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: checkout-success-rate
namespace: production
spec:
args:
- name: canary-hash
metrics:
- name: success-rate
interval: 1m
count: 5
failureLimit: 1
successCondition: result[0] >= 0.99
failureCondition: result[0] < 0.95
provider:
prometheus:
address: http://prometheus.monitoring.svc:9090
query: |
(
sum(rate(http_requests_total{app="checkout",revision="{{args.canary-hash}}",status!~"5.."}[5m])) or vector(0)
)
/
clamp_min(
sum(rate(http_requests_total{app="checkout",revision="{{args.canary-hash}}"}[5m])) or vector(0),
1e-9
)
- name: request-volume
interval: 1m
count: 5
failureLimit: 1
successCondition: result[0] >= 100
failureCondition: result[0] < 100
provider:
prometheus:
address: http://prometheus.monitoring.svc:9090
query: |
sum(increase(http_requests_total{app="checkout",revision="{{args.canary-hash}}"}[5m])) or vector(0)
This example assumes the Prometheus time series labels and the status value exist as shown; many environments use a different metric schema. The or vector(0) fallback and denominator clamp intentionally turn missing or zero-volume candidate data into a zero success rate, while the separate request-volume metric enforces a sample floor. The threshold of 100 requests per five-minute window is illustrative, not portable; tune it to expected canary traffic and service risk. This is a fail-closed policy: repeated no-data measurements fail the analysis rather than silently passing. failureLimit controls tolerated failed measurements; it does not replace service-level alerting. If monitoring outages should pause for human review instead of count as candidate failure, use a separately tested provider-error policy. Validate the query and no-data behavior against real series before enabling the gate.
Analysis can run inline at a step, in the background while a canary progresses, before a blue-green traffic switch, or after promotion depending on strategy. Pre-promotion analysis can prevent a switch until the candidate passes a check. Post-promotion analysis detects regressions after real traffic is sent and can trigger abort behavior, but the effect on user requests depends on how quickly the controller and traffic provider reconcile. Decide whether the experiment’s goal requires pre-traffic validation, low-weight traffic, or both.
Connect analysis to rollout steps
Make pauses and weights reflect risk. A low initial weight limits exposure but can produce too little traffic for statistical evidence. A step should remain long enough to collect the intended metric window after routing changes have propagated. A manual pause can provide a deliberate human checkpoint, but should include an owner and a timeout policy so a stale release does not stay indefinitely in an ambiguous state.
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: checkout
namespace: production
spec:
replicas: 8
selector:
matchLabels:
app: checkout
template:
metadata:
labels:
app: checkout
spec:
containers:
- name: app
image: registry.example.invalid/checkout@sha256:REPLACE_WITH_APPROVED_DIGEST
ports:
- name: http
containerPort: 8080
strategy:
canary:
stableService: checkout-stable
canaryService: checkout-canary
steps:
- setWeight: 10
- pause:
duration: 5m
- analysis:
templates:
- templateName: checkout-success-rate
args:
- name: canary-hash
valueFrom:
podTemplateHashValue: Latest
- setWeight: 50
- pause: {}
The manifest is a representative Rollout fragment, not a complete ingress or Service configuration. The image is a placeholder and must be replaced with a verified immutable digest. Define Services and the traffic-routing integration supported by your cluster, and confirm their selectors point to the Rollout’s Pods. Without routing integration, weight behavior can be only an approximation based on replica scaling. Verify the installed Argo Rollouts CRD and controller version before applying the manifest.
Interpret outcomes deliberately
An AnalysisRun can finish as Successful, Failed, Inconclusive, or Error, depending on metric results and provider behavior. An Inconclusive run pauses the Rollout at its current step for manual intervention; resuming it without resolving missing or ambiguous evidence can promote a release without a valid gate. Conversely, treating every transient monitoring outage as a product regression can block releases unnecessarily. Define and test these outcome semantics for the installed Argo Rollouts version. Set bounded measurement counts, an interval, and a timeout that align with the Prometheus scrape interval, query range, and traffic volume.
If an analysis fails, the controller can mark the rollout degraded and abort a canary. An abort changes the desired traffic split and stable/canary state, but it cannot reverse external side effects already created by the application, schema migration, message publication, or data transformation. Rollback is effective only if the previous revision remains runnable and backward-compatible with state changes introduced by the candidate.
Use pre-promotion analysis for checks that can be evaluated against the preview service without exposing ordinary traffic, such as smoke tests, synthetic requests, or compatibility checks. Use background analysis for an ongoing canary signal that must be sampled while the revision receives traffic. Use post-promotion analysis only when the release process tolerates a short exposure window and the traffic rollback path has been exercised. Avoid duplicating the same expensive query at every stage without considering Prometheus load.
Protect query quality and telemetry cardinality
The metric must distinguish candidate traffic from stable traffic. If both versions use the same labels or the query aggregates them together, the canary’s error rate can be hidden by the stable fleet. Add a revision or ReplicaSet label through instrumentation or routing telemetry, and ensure the label survives the metrics pipeline. Avoid unique request IDs, user IDs, or unbounded tenant labels that create high-cardinality series.
A rate query needs a window that contains enough samples for the scrape interval. A five-minute range with a one-minute scrape may have sparse or delayed samples during startup; validate at the same range and label set the analysis will use. Account for Prometheus lookback, delayed remote write, downsampling, and clock skew. If the canary sends only a small fraction of production requests, a single error can create a high ratio; combine rate with request-count evidence or choose a meaningful minimum sample count.
Test the AnalysisTemplate in a non-production environment with known good, known failing, no-data, and Prometheus-unavailable cases. Confirm the AnalysisRun status and the resulting Rollout state. Store the query and thresholds in code review, and require an owner for exceptions. Dashboards are useful for humans, but the automated query should be versioned and deterministic rather than copied manually from a dashboard panel.
Operate promotion and abort as change controls
Use the kubectl argo rollouts plugin or controller status to observe steps, AnalysisRuns, pause conditions, ReplicaSets, and events. The plugin makes state changes visible but does not replace cluster RBAC or a documented approval process. Restrict who can promote or abort production Rollouts, and keep deployment manifests protected. If a rollout is paused, determine whether it is waiting for a manual action, an analysis result, a provider, or a controller retry before issuing a command.
For each application, define the stable revision retention window, max surge and unavailable policy, traffic-provider health expectations, metric thresholds, abort owner, and rollback compatibility. Test a failed analysis and a manual abort before adopting the pattern for critical services. Record analysis query versions alongside application releases so responders can identify whether a changed threshold or label selector altered the gate.
Progressive delivery is a measurement system wrapped around a deployment controller. Correct routing, representative metrics, explicit no-data semantics, and a tested abort path matter more than the number of canary steps. Keep the canary hypothesis narrow enough to be measurable, and make every automated decision explainable from the recorded AnalysisRun.
Related:
- How to Set Up Canary Analysis with Automated Rollback
- How to Order Argo CD Deployments with Sync Phases, Waves, and Hooks
Sources: