Prometheus Federation: Hierarchical Aggregation, Selectors, and Failure Boundaries
Federate selected current Prometheus series for global views without mistaking federation for replication, remote write, or a complete historical data copy.
Prometheus federation is a pull-based way for one Prometheus server to scrape selected time series from another Prometheus server. It is useful when a global view needs a bounded set of pre-aggregated metrics from several local monitoring systems, or when a service needs a selected set of related metrics from another domain. It is not a replica of the source server, a bulk transfer of its historical samples, or a substitute for remote write.
That boundary should drive the design. The source keeps its detailed local data and remains the place for local drill-down. The destination scrapes the source’s /federate endpoint and stores the selected samples it sees over time. A broad selector can still move a large number of series, so federation needs an explicit metric contract, cardinality budget, source identity strategy, and observable scrape path.
Choose federation for a bounded global view
In hierarchical federation, regional or cluster Prometheus servers retain detailed instance-level series. Higher-level servers pull a smaller set of recorded aggregates, such as request rate or error ratio by service and region. This provides a global view without making the top-level server scrape every application and machine endpoint directly.
Cross-service federation is another fit: one server may pull a carefully selected set of infrastructure or dependency metrics from a separate Prometheus server so its operators can query related signals together. In both designs, choose series based on the questions the destination must answer. Federating every raw series defeats the purpose of a bounded aggregation layer and can recreate the same storage and cardinality pressure at another tier.
Federation differs from several nearby mechanisms:
| Mechanism | Data path | Appropriate expectation |
|---|---|---|
| Federation | Destination scrapes selected series from source /federate |
Selected current values are sampled into the destination’s own TSDB over time |
| Remote write | Source queues ingested samples and sends them to a receiver | Streaming samples to a remote-compatible storage endpoint |
| Remote read | A querying server requests raw series data and time ranges from another endpoint | Query-time access to remote raw data, with the querying Prometheus still evaluating PromQL |
| Direct scraping | A Prometheus server scrapes exporters or application endpoints | Local collection from instrumented targets |
Use the mechanism that matches retention, query, and failure requirements. If the destination must retain every source sample at the source’s resolution, federation’s selector and scrape interval are not a lossless replication plan. If users need detailed historical queries from the source, keep a source link or use an appropriate remote-storage/query architecture instead of assuming the global server contains the full history.
Select series with match[]
The source exposes selected time series through /federate. Each request must contain at least one match[] parameter, and every value is an instant-vector selector. Multiple parameters select the union of their matches. A selector can match a recorded series by metric name or filter series by labels:
/federate?match[]={job="prometheus"}
/federate?match[]={__name__=~"service_region:.*"}
Prefer a narrow family of recording rules whose names and label dimensions are intentionally designed for federation. For example, a local Prometheus can record request rate by service and region, then expose only those aggregates to the global tier. Include only labels the global dashboards and alerts require. A label such as request ID, user ID, or raw URL can turn one aggregate into an uncontrolled set of time series and erase the intended storage reduction.
Multiple selectors form a union, not an instruction to copy a historical time range. Review the effective selector against the source before rollout. A broad metric-name regex can match new series added months later, silently expanding the federation contract. Prefer explicit prefixes and labels that make additions deliberate, and review selector changes alongside recording-rule changes.
Configure a destination scrape job
The destination uses a normal scrape_config with the source Prometheus endpoints, a /federate path, the desired selectors, and honor_labels: true so source labels are preserved on conflicts. This example follows the documented pattern and uses a small, deliberate selector set:
scrape_configs:
- job_name: federate-region-a
scrape_interval: 30s
scrape_timeout: 10s
honor_labels: true
metrics_path: /federate
params:
'match[]':
- '{job="prometheus"}'
- '{__name__=~"service_region:.*"}'
static_configs:
- targets:
- prometheus-a.internal.example:9090
Use a distinct, stable job_name for each source or source group where the distinction helps operations. With honor_labels: true, labels in the federated series take precedence over conflicting scrape-side labels, which preserves the source series identity but means the destination cannot assume its own job or instance values will replace them. Test the final label set in the destination before relying on it in dashboards or alert rules.
Use a stable source or cluster label so identical metric names from different sources remain distinguishable. Establish where that label is added and confirm it is present in the actual federated output. If multiple source servers expose overlapping series with identical labels, the destination may encounter duplicate series identities or ambiguous operational views. High-availability source pairs need a deliberate strategy for preserving source identity or deduplicating downstream; federation itself does not turn two servers into one deduplicated data source.
An appropriate scrape interval is part of the contract. If a local rule is evaluated every minute and the global server scrapes every five seconds, most federation requests retrieve the same most recent recorded value and add cost without creating new information. If the destination scrapes too slowly for the alert or dashboard freshness objective, it increases the delay before the value is visible there. Coordinate the source rule interval, destination scrape interval, timeout, and service objective rather than copying the default interval blindly.
Keep aggregate queries semantically safe
Federated series should answer a stable query at the destination. A local recording rule that sums by service and region may be suitable for fleet-level capacity views; a raw per-instance latency histogram or counter with high-cardinality labels may be better kept local unless the global use case genuinely needs that detail.
Preserve enough labels to avoid accidental cross-source aggregation. If a global query sums requests:rate5m without grouping or filtering by region, it may be correct for a worldwide total but wrong for a region-specific SLO. Conversely, retaining every low-level label “just in case” may grow the global TSDB. Write example PromQL for intended destination queries before approving the selector list, then check the resulting output labels and cardinality against those queries.
Be explicit about missing and stale sources. A failed federation scrape makes the destination’s federated target unhealthy; it does not mean that the source’s own alerting and local measurements stopped. Destination users should be able to tell the difference between “service metric is zero,” “the local recording rule emitted no series,” and “the source could not be scraped.” Alert on federation target health separately from the application SLO and provide a link to the source server’s local view.
Native histograms need explicit handling
If federation includes native histograms, the scraping Prometheus configuration must set scrape_native_histograms: true. When scrape_protocols is not explicitly configured, this makes Prometheus prefer the protobuf exposition format. An explicit protocol list can override that default; because Prometheus protobuf is currently the only format that carries native histogram samples, verify both ingestion and protocol negotiation on the exact releases at each end before enabling this in a mixed-version fleet.
The current Prometheus federation documentation also calls out an edge case where the same metric name appears with mixed sample types. The federation payload can contain multiple metric families with that name and different types; float samples are federated as untyped, while histogram samples retain their histogram type. Prometheus can ingest that payload, although it is technically outside the protobuf exposition format’s usual rules. Avoid treating a successful scrape as proof that downstream tools, remote systems, and queries handle every mixed-type case as intended. Test representative histogram queries and dashboards before widening the selector.
If the global view needs only a histogram-derived quantile or a classic aggregate, decide whether to federate a suitable recording rule instead of transporting native histograms. Do not collapse a histogram to an average and then describe it as a tail-latency percentile; that changes the statistical meaning of the signal.
Budget load and protect the scrape path
Estimate the number of series matched per source and the number of destinations that will scrape it. The request cost is paid at the source endpoint, and the destination then ingests the selected results at its own scrape cadence. A regex that is safe for one global server may become expensive if copied to several independent consumers.
Start from a test environment or a limited selector. Measure source query and scrape duration, selected series count, destination ingestion rate, and TSDB growth. Expand only when those measurements fit the storage and latency budget. Re-evaluate the budget when a recording rule gains labels, a new source is added, or a regex begins matching a new metric family.
The /federate endpoint is an HTTP scrape endpoint. Restrict reachability to intended Prometheus clients and configure the server’s supported HTTPS or authentication path according to the current Prometheus deployment guidance. Do not publish an unauthenticated source endpoint as a casual Internet-facing metrics API. Keep credentials in the deployment’s supported secret mechanism and test that both successful and unauthorized requests behave as intended.
For each source, monitor the federation scrape’s up series and scrape duration, and distinguish those from application-level metrics. A successful connection can still return a selector that matches no series. Add destination-side checks that assert expected metric families and labels exist, while keeping those checks tolerant of a legitimate zero value. This catches typoed selectors, renamed recording rules, and accidental loss of a region label before dashboards quietly become incomplete.
Roll out, change, and recover deliberately
Treat the selector set and label schema as an interface. Document the source owner, destination owner, metric names, labels, scrape interval, expected series budget, retention requirements, and deprecation process. Add a new selector in a canary destination or a low-impact job first, then compare results with direct source queries over a representative workload window.
When changing a recording rule, check both sides of the interface. A renamed recording metric can leave the destination up but ingesting none of the data it expects. A new label can multiply series and change the destination’s query grouping. A source URL change can turn a healthy data path into a DNS, TLS, or routing failure. Review federation changes in the same release as the producer rule changes where possible.
For recovery, first determine which layer failed: source Prometheus health, rule evaluation, /federate selector, network/authentication, destination scrape, or destination ingestion. Query the source locally to verify the series, request the /federate endpoint with the exact match[] selectors, inspect the destination target health and scrape errors, then verify the ingested label set. Avoid widening selectors or changing all intervals while the failing layer is unknown.
Production acceptance checklist
- The architecture states why federation is preferable to remote write, remote read, or direct scraping.
- The
/federateselectors are narrow, documented, and based on stable recorded series where appropriate. - The destination uses
honor_labels: trueand the resulting source and target label precedence has been checked. - A stable source identity is preserved when multiple clusters or HA servers publish overlapping metrics.
- Source evaluation cadence, destination scrape interval, timeout, and freshness objective are compatible.
- Matched series count, source request duration, destination ingestion rate, and TSDB growth are measured.
- Missing series, stale source, failed scrape, and real zero values are distinguishable in alerts and dashboards.
- Native histogram behavior is explicitly configured and tested when histograms are included.
- The endpoint is reachable only by intended scrapers and its transport/authentication behavior is tested.
- Runbooks link destination metrics back to the source’s detailed view.
Federation works best as a deliberately small interface between Prometheus servers: the source retains detailed observability, while the destination receives only the signals it needs for cross-region or cross-service decisions. Keep that interface explicit, measure its cost, and verify the samples that arrive rather than assuming that a green target means the global view is complete.
Related:
- Prometheus Remote Write: Backpressure, WAL Recovery, and Delivery
- Prometheus Scrape Relabeling: Target, Sample, and Remote-Write Stages
Sources: