Prometheus Scrape Relabeling: Target, Sample, and Remote-Write Stages
Trace Prometheus target, sample, and remote-write relabeling stages, label collisions, cardinality effects, and safe validation techniques.
Prometheus exposes one relabeling language in several places, but the stage in which a rule runs determines what it can see and which data it can affect. relabel_configs changes a discovered target before Prometheus scrapes it. metric_relabel_configs transforms individual scraped samples at the last step before local ingestion. write_relabel_configs filters or rewrites the sample stream headed to one remote-write endpoint, after external labels are applied. Confusing those stages leads to silent target loss, ineffective filters, or different local and remote datasets.
Relabeling should be treated as part of the metrics schema. A rule can change target routing, URL path, scrape protocol settings, the labels that identify every series from a target, or the samples that remain at a destination. Review it with the same care as an instrumentation change: a label change can split historical series, merge previously distinct series, change alert behavior, and alter storage cost.
The three stages have different inputs
The high-level flow is:
- Service discovery or static configuration produces targets and labels.
- Target relabeling runs in order. A
keepordropaction can remove a target before the scrape; special labels can change its address, scheme, path, parameters, or per-target scrape options. - Prometheus scrapes each retained target. It combines labels exposed by the target with labels attached by Prometheus;
honor_labelsdefines how conflicts are resolved. - Metric relabeling runs on scraped samples, after target labels have been applied and immediately before local ingestion. It cannot prevent the HTTP request or remove automatically generated series such as
up. - For an external destination,
external_labelsare added when a series does not already contain those labels. Each remote-write configuration can then apply its ownwrite_relabel_configsbefore sending.
Target relabeling has access to labels such as __address__, __scheme__, __metrics_path__, __param_<name>, and discovery-specific __meta_* labels. Labels beginning with __ are removed after target relabeling unless a rule copies their value to an ordinary label. The __tmp prefix is reserved for temporary working labels. This makes target relabeling appropriate for selecting, routing, or annotating scrape targets, not for matching metric names that have not been scraped yet.
Target selection and endpoint rewriting
This example accepts only annotated Pods in the payments namespace, changes the scrape path when an annotation is present, and copies two bounded pieces of Kubernetes metadata into ordinary target labels. Kubernetes service-discovery metadata varies by role and version; verify the labels emitted by the exact discovery configuration you use.
scrape_configs:
- job_name: payments-pods
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_namespace]
action: keep
regex: payments
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: "true"
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path]
action: replace
regex: "(.+)"
replacement: "$1"
target_label: __metrics_path__
- source_labels: [__meta_kubernetes_pod_label_app]
action: replace
regex: "(.+)"
replacement: "$1"
target_label: app
- source_labels: [__meta_kubernetes_pod_label_environment]
action: replace
regex: "(.+)"
replacement: "$1"
target_label: environment
Rules execute in listed order. The keep filters mean missing namespace or scrape annotation values do not match and therefore the target is removed. A replace rule with a nonmatching regex leaves the target unchanged, so a missing path annotation keeps the configured default metrics path. Here the regex is anchored by Prometheus, and (.+) requires a nonempty value. Review that fallback deliberately; an accidental keep rule for an optional label can eliminate every target.
source_labels values are concatenated with the configured separator (semicolon by default); a missing label contributes an empty string. Regexes use RE2 syntax and are anchored at both ends unless the expression explicitly includes .*. For example, regex: payments matches exactly payments, not payments-staging. A replace rule can use capture groups to construct a label or special setting. Use temporary labels with the __tmp prefix when a multi-step rewrite needs an intermediate value, then confirm that temporary state is not copied to final target labels.
Rewriting __address__, __scheme__, __metrics_path__, or __param_* changes how Prometheus makes the scrape request. It is easy to accidentally redirect traffic to an unintended endpoint, scrape a privileged management path, or create a target with a missing port. Keep the input set narrow, preserve TLS and authentication settings, and inspect the final active target URL in the Prometheus Targets page after a staged reload.
Target labels become part of every series identity
The target’s job label starts from job_name; instance defaults to the final __address__ unless already set during target relabeling. Ordinary target labels are attached to scraped samples. If the exporter also emits a conflicting label, the default honor_labels: false keeps the Prometheus-side value and renames the exporter’s conflicting value to exported_<name>. With honor_labels: true, Prometheus preserves the scraped value and ignores the conflicting server-side value. Enable honor_labels only when that is the intended contract, such as some federation or Pushgateway cases.
Each unique metric name and complete label set identifies a time series. Adding a pod_uid, request path, or ephemeral endpoint label can multiply the number of series; removing a distinguishing label can collapse different streams onto the same identity. Normalizing labels can reduce churn, but only if the aggregation still represents the same measurement. A label transformation is not a retroactive rename: old and new label sets are distinct over time, so dashboards and rules may see separate histories.
Do not indiscriminately labeldrop job, instance, cluster, or tenant dimensions to make series count look smaller. If label removal maps samples onto the same series and timestamp but their values differ, the current Prometheus TSDB append path can reject the conflicting sample as a duplicate-timestamp error; an exact repeat with the same value is accepted by that path. Treat this as implementation behavior to verify against the deployed Prometheus release. Even when timestamps differ, distinct sources may be merged into a misleading series. Before collapsing dimensions, identify the owner and intended aggregation semantics, query for collisions in a staging environment, and update dashboards, recording rules, and alert selectors together.
Metric relabeling is the final pre-ingestion filter
Metric relabeling works on a scraped sample’s label set, including __name__. It is useful when the exporter cannot be changed quickly or one scrape job should omit a specific family. For example:
metric_relabel_configs:
- source_labels: [route]
action: replace
regex: "/orders/[0-9]+"
replacement: "/orders/:id"
target_label: route
- source_labels: [__name__]
action: drop
regex: "debug_payload_bytes"
The first rule normalizes only route values matching the complete expression. The second drops the named sample family from local ingestion for this scrape job. Because target metadata such as environment has already been copied into ordinary labels, metric relabel rules can also use it to make a narrowly scoped keep/drop decision. Prefer explicit source_labels, anchored expressions, and a named job boundary over a broad labelmap that imports every discovered label into every series.
This stage has important limits. It runs after Prometheus fetched and parsed the response, so dropping a sample does not save network bandwidth, exporter CPU, or the memory required to parse the exposition. It also does not apply to automatically generated time series such as up; to stop scraping an endpoint, filter the target earlier. Per-scrape sample_limit, label_limit, and label name/value length limits are checked after metric relabeling, and exceeding them can fail the entire scrape rather than preserve the subset that fits. Use limits as circuit breakers, not as a substitute for a valid metric contract.
Remote-write relabeling is destination-specific
External labels identify the Prometheus instance or cluster in communication with external systems. They are applied only when a series lacks a label with the same name. A remote-write block can then filter outgoing samples after those external labels have been applied. This changes what that endpoint receives; it does not erase the series from the local TSDB.
global:
external_labels:
cluster: prod-eu-1
remote_write:
- name: long-term-prod
url: https://metrics.example.net/api/v1/write
write_relabel_configs:
- source_labels: [environment]
action: keep
regex: production
- source_labels: [__name__]
action: drop
regex: "debug_.*"
This remote endpoint receives only series with an environment="production" label after normal scrape processing, excluding metric names with the debug_ prefix. If no series has an environment label, the keep rule drops those series from this remote stream. Another remote_write configuration can apply a different rule set. Confirm the desired local-versus-remote retention explicitly before using write relabeling as a cost control.
Do not attach a Prometheus replica label through global external_labels and then assume honor_labels will override it; external-label conflict behavior is separate. For highly available Prometheus pairs, choose stable cluster identity and replica identity intentionally, and make the remote backend’s deduplication contract explicit. A value that changes on restart or rollout can fragment remote series even when local scrape labels appear stable.
Validate from discovery through the destination
Run promtool check config prometheus.yml against the version you deploy before reload. A syntactically valid configuration still may select zero targets, map an annotation incorrectly, or alter series identities unexpectedly. Compare active and dropped targets in the Prometheus UI or API, inspect the final scrape URL and labels, and verify representative metrics before and after the change.
For a rollout, estimate the expected target count and series change first. On a staging instance, compare the target’s scraped exposition with local query results, confirm up behavior separately from metric drops, and inspect exported_* labels if the exporter and Prometheus attach the same name. For a remote-write filter, query the local TSDB and remote backend independently so that a successful local scrape is not mistaken for successful remote delivery.
After reload, check scrape errors, sample-limit failures, rejected samples, active-series count, remote-write queue health, and alerts whose selectors depend on changed labels. Keep a rollback copy of the known-good relabel chain. When a rule drops data, record the exact stage and destination: “the metric is gone” is not enough to determine whether discovery, scraping, local ingestion, or only remote forwarding was affected.
Relabeling is powerful because one ordered language operates at several boundaries. Production safety comes from naming those boundaries, testing the labels visible at each one, and treating changes to final labels as changes to the schema consumers query.
Related:
- Fixing Prometheus Cardinality Explosions Before They Exhaust Memory and Storage
- Prometheus Alert Rules: Evaluation, Pending State, and Notification Boundaries
Sources: