Prometheus Rule Tests: promtool Fixtures and Alert Coverage
Test PromQL and alert-rule behavior deterministically with promtool, including pending intervals, labels, missing series, group order, and CI validation.
Prometheus rule tests turn alert and recording-rule behavior into a repeatable check. With promtool test rules, a test supplies synthetic time series, evaluates the configured rules at known times, and compares the result with an expected alert or PromQL sample. This catches mistakes in expressions, label sets, annotations, and for timing before a rule reaches the production rule manager.
Rule tests have a deliberately limited boundary. They validate rules against supplied samples; they do not scrape a target, prove service discovery is correct, load a real Prometheus server configuration, or verify Alertmanager routing and receiver delivery. Treat them as deterministic unit tests in a broader pipeline, not as proof that an on-call engineer will receive a page.
Keep rule files and test fixtures separate
A rule test file points at one or more actual rule files, sets an evaluation interval, and defines test groups. Each group provides synthetic input series and assertions for alerting rules or PromQL expressions. The rule file is the production logic; the test file is the fixture and expected behavior:
# rules/instance-health.rules.yml
groups:
- name: instance-health
rules:
- alert: InstanceDown
expr: up == 0
for: 2m
labels:
severity: page
annotations:
summary: "{{ $labels.job }} instance {{ $labels.instance }} is down"
# tests/instance-health.test.yml
rule_files:
- ../rules/instance-health.rules.yml
evaluation_interval: 1m
tests:
- name: sustained outage becomes a page
interval: 1m
input_series:
- series: 'up{job="checkout",instance="api-0"}'
values: '0 0 0'
alert_rule_test:
- eval_time: 1m
alertname: InstanceDown
exp_alerts: []
- eval_time: 2m
alertname: InstanceDown
exp_alerts:
- exp_labels:
job: checkout
instance: api-0
severity: page
exp_annotations:
summary: "checkout instance api-0 is down"
The first evaluation starts the alert’s pending period. At one minute, the expression is still true but the two-minute for interval has not elapsed, so no firing alert is expected. At two minutes, the alert should fire with the expected labels and rendered annotation. This test makes the timing and identity contract explicit instead of relying on someone to eyeball the rule YAML.
Keep fixture files near the rules they exercise, but do not accidentally include test-only metadata in the production rules directory loaded by Prometheus. A test file can reference multiple rule files and supports globs; make that relationship clear in the repository layout and CI command.
Build representative time-series fixtures
Each input_series entry declares a full Prometheus series selector and a sequence of values. The interval determines the spacing between samples for that test group; if omitted, the test uses the evaluation interval. Start with the smallest fixture that distinguishes correct behavior from a likely regression.
Test both sides of important boundaries:
- A condition that remains true long enough to fire, and a condition that resolves before its
forduration expires. - A series that disappears or changes labels, so the rule’s behavior for absent data is intentional.
- Counter resets or gaps where the expression uses
rate()orincrease(). - A zero-traffic or low-sample case where a ratio could divide by zero or misrepresent a service.
- Expected output labels after aggregation, joins, and rule-level labels are applied.
- Annotation templates with the labels and values that the production rule is expected to expose.
Prometheus supports compact value notation such as 1+0x5 to represent repeated samples. It is useful for long fixtures, but expanded values can be easier for reviewers to interpret. Prefer clarity over compression and comment any non-obvious sequence. If a test depends on calendar functions or time(), set start_timestamp explicitly so the clock context is deterministic; by default, rule test time begins at the Unix epoch.
Avoid making the fixture mirror only the exact current production data. Include edge cases that protect the rule’s operational promise, such as an absent target, label changes, no traffic, or a brief interruption. A test can pass while still being unrepresentative if all its series are perfectly regular and every input label is always present.
Assert the PromQL result separately
promql_expr_test checks a query or recording-rule expression at a selected eval_time. It is useful for validating aggregations, thresholds, label grouping, and conversion functions independently of the alert state machine:
promql_expr_test:
- expr: sum by (job) (rate(http_requests_total[2m]))
eval_time: 3m
exp_samples:
- labels: '{job="checkout"}'
value: 12
The expected label set must match the expression’s actual output, including whether the metric name is retained. If a rule adds or overwrites labels, test that behavior through the alerting-rule assertion or a query of the resulting recording-rule series. Avoid using one broad expected result that would let an accidental extra label or duplicate series pass unnoticed.
Use fuzzy_compare: true only when the expected and actual difference is a tiny floating-point representation detail. It slightly weakens numeric comparison; it is not a way to hide a wrong rate window, wrong aggregation, or materially inaccurate result. State the tolerated precision in the test’s comments when the tolerance affects a meaningful threshold.
For dependent recording rules, remember that rules in a group are evaluated sequentially, but group ordering is not generally guaranteed unless the test defines group_eval_order. Include the relevant group names in the expected order when the behavior depends on one group’s output being available to another at the same evaluation time. Prefer putting rules that intentionally depend on each other in the same group when that is the intended production design.
Cover alert identity and lifecycle, not just the expression
An alert expression can be correct while the alert is still operationally wrong. The output label set defines alert identity; adding a high-cardinality label can create a separate alert for every value, while removing a distinguishing label can merge independent incidents. Assert the final expected labels and annotations, including rule-defined severity and the dimensions an operator needs to identify the affected service.
For for and keep_firing_for, test the lifecycle across multiple evaluations. A single evaluation cannot prove that an alert remains pending for the intended duration, becomes firing at the correct point, resolves when the expression stops matching, or stays firing for its configured hold period. Test a short interruption as well as a sustained outage. If the rule has a for duration longer than the test data, the expected alert may still be pending rather than firing.
Tests for missing data should be explicit. In PromQL, an absent series is not automatically equivalent to a numeric zero; expressions such as up == 0, absent(), vector matching, and or have different behavior. Provide a fixture with no matching series and state whether the desired behavior is to alert, remain silent, or rely on a separate target-discovery alert. Do not fabricate a zero-valued series when the production failure mode is that the series disappears entirely.
Recording-rule tests should verify the resulting metric name, labels, and numeric value. If a recording rule is intended to reduce query cost, promtool verifies its logic but does not benchmark production query latency, rule-group duration, or cardinality. Measure those separately after deployment and monitor missed rule evaluations.
Run syntax checks and tests in CI
Use a promtool binary from the same Prometheus release family as production. This catches differences in supported rule syntax and avoids validating with a parser that is newer than the server that will load the rules:
promtool check rules rules/*.yml
promtool test rules tests/*.test.yml
promtool check rules confirms rule-file syntax and performs the available checks; promtool test rules evaluates the fixtures and expected results. Run both on every pull request that changes a rule or its test. Make the CI job fail on a non-zero exit code and preserve the command output so a mismatch is visible in review.
Syntax and unit-test success do not verify that the production Prometheus config includes the intended file under rule_files. Also check the full server configuration with the release-matched tool, inspect the loaded rules after rollout, and check for evaluation errors and missed iterations. If rules are reloaded dynamically, reload only a fully validated set of files; Prometheus applies the new rules only when all configured rule files are well-formed.
Finally, test the delivery path separately. Prometheus can show an alert as firing while Alertmanager is unreachable, a route does not match, a silence suppresses it, or a notification receiver rejects the request. Use a controlled test alert and verify the resulting notification in a non-paging destination. Keep that integration test distinct from promtool unit tests so failures identify the layer that needs attention.
Production acceptance checklist
- Every changed alert or recording rule has a test fixture that exercises its intended behavior.
- The fixture covers boundary timing, expected labels and annotations, and relevant missing-data or reset behavior.
group_eval_orderis explicit when tests rely on cross-group rule dependencies.- Floating-point tolerance is narrow and justified rather than used to conceal semantic differences.
- CI runs both version-matched
promtool check rulesandpromtool test ruleson every relevant change. - The server configuration actually loads the validated rule files, and post-deploy rule evaluation is monitored.
- Alertmanager routing, inhibition, silences, and receiver delivery are tested independently of rule-unit results.
- The team can explain what the tests do not cover: scrapes, discovery, live server reload, long-range data, query performance, and notification delivery.
Rule tests are most valuable when they encode the operational contract of a rule: which input patterns should alert, after how long, with which identity and useful context. A passing promtool run is strong evidence for that deterministic logic, while a separate integration check proves the rest of the monitoring pipeline can turn it into an actionable notification.
Related:
- Prometheus Alert Rules: Evaluation, Pending State, and Notification Boundaries
- Alertmanager High Availability: Peering, Deduplication, and Delivery
Sources: