Prometheus Exemplars: Trace Correlation, Storage, and Delivery
Connect Prometheus measurements to distributed traces with exemplars, managing experimental storage, remote-write support, queries, and rollout risk.
Prometheus exemplars attach a small set of labels and a value to a metric sample, providing a reference to related data outside the metric series. A trace identifier is a common use: an engineer sees a latency distribution move, selects an exemplar near the interesting measurement, and follows its trace in the tracing system. This connects aggregate metrics to individual requests without putting a unique request or trace ID on every time series.
Exemplars are not full traces, a replacement for traces, or an automatic end-to-end integration. The application must emit exemplar data, Prometheus must accept and store it, any remote-write receiver must preserve it, and the query or visualization layer must resolve the identifiers to a trace. Current Prometheus documentation still classifies exemplar storage and its query endpoint as experimental. Treat enablement, API shape, receiver compatibility, and upgrade behavior as release-specific until the project marks them stable.
What an exemplar contains
An exemplar is attached to a measured sample but does not become part of that series’ label set. It can carry a bounded label set such as trace_id and span_id, the exemplar value, and a timestamp. OpenMetrics 2.0 recommends the names trace_id and span_id so that consumers can correlate traces consistently. Confirm which exposition format and exemplar support your instrumentation library and scrape target actually provide.
The example below shows the shape of an OpenMetrics sample with an exemplar. The identifiers are placeholders and must be replaced by valid identifiers from a real trace backend:
http_request_duration_seconds_bucket{le="0.5",service="checkout"} 42 # {trace_id="4bf92f3577b34da6a3ce929d0e0e4736",span_id="00f067aa0ba902b7"} 0.42 1730000000.125
The metric labels still describe a bounded service, route, or status dimension. The exemplar points to one representative observation. Do not move trace_id, request ID, user ID, or an unbounded URL into metric labels to imitate exemplar behavior: those labels would create separate time series and raise cardinality for every unique value.
For latency metrics, exemplars are especially useful because they let a team inspect a concrete request associated with a measured bucket or histogram observation. Sampling remains selective. A metric continues to represent the full population according to its instrumentation and scrape behavior; exemplars are sparse breadcrumbs, not a statistically complete list of slow requests. Use tracing sampling and exemplars together, and keep a separate logging or trace-search workflow for cases where the relevant request did not produce an exemplar.
Enable local Prometheus storage carefully
Prometheus currently requires the exemplar-storage feature flag to ingest and store scraped exemplars. The configuration below enables the feature and sets the circular buffer capacity. Use the flags and field names supported by the exact server version in production:
prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--enable-feature=exemplar-storage
# prometheus.yml
storage:
exemplars:
max_exemplars: 50000
The exemplar store is a fixed-size in-memory circular buffer shared across series. The documented default is 100,000 exemplars. When it fills, older entries roll out as new exemplars arrive; the limit is a capacity, not a promise to keep a fixed duration. The same capacity can cover seconds in a very high-volume server or much longer in a quieter one. Estimate a safe value from observed exemplar rate, memory use, and the time range operators need, then verify it under real load. Prometheus documentation estimates roughly 100 bytes for an exemplar containing only a trace ID, but treat that as a rough planning value rather than a total process-memory guarantee.
Enabling storage also appends exemplars to the write-ahead log for its local persistence window. This does not turn exemplar storage into a long-retention trace database. Plan capacity and retention independently from metric blocks and from the tracing backend. A server restart, WAL lifecycle, and buffer-size change should be included in acceptance tests for the deployed release.
Preserve exemplars across remote write only when supported
Local scrape ingestion and remote forwarding are separate decisions. The Prometheus remote_write configuration has a send_exemplars option that defaults to false; exemplar storage must already be enabled for Prometheus to scrape exemplars in the first place.
remote_write:
- url: https://metrics.example.invalid/api/v1/write
send_exemplars: true
Enable this only after confirming that the selected message format, sender version, intermediary, and receiver all support exemplar data. A remote-write success for ordinary samples does not prove that exemplars survived the path. Check the receiver’s documentation and test its query or trace-link workflow with a known exemplar before enabling the feature fleet-wide. Some systems accept the request but may have different storage limits or query interfaces.
This introduces another observability boundary: the local Prometheus may have the exemplar while the remote backend does not. Decide which query system is authoritative for dashboards and traces, verify that it can find the exemplar in the relevant time range, and alert on remote-write failures separately from scrape failures. When send_exemplars is off, that is a deliberate data-path choice, not a bug in scraping.
Query exemplars and prove the path
The Prometheus HTTP API documents /api/v1/query_exemplars for retrieving exemplars associated with a PromQL expression over a time range. The endpoint is currently experimental. Use the API only from a client compatible with the deployed server, and avoid binding automation to response details not guaranteed by that release:
curl -G 'http://localhost:9090/api/v1/query_exemplars' \
--data-urlencode 'query=http_request_duration_seconds_bucket{job="checkout"}' \
--data-urlencode 'start=2026-10-03T12:00:00Z' \
--data-urlencode 'end=2026-10-03T12:10:00Z'
The selector must match a series that can carry the exemplar, and the requested interval must include its timestamp. The API result identifies the series labels and returns the exemplar labels, value, and timestamp. An empty result does not by itself prove an instrumentation defect: the time range, query, protocol negotiation, feature flag, sample format, circular-buffer eviction, or remote receiver could be responsible.
Run a controlled validation with one instrumented endpoint and a known trace:
- Confirm the application’s metrics endpoint emits an exemplar using the format negotiated by the scrape.
- Check Prometheus target health and the loaded server feature flags; target health alone does not show that an exemplar was parsed.
- Query the exemplar API with a narrow selector and a time range around the known observation.
- Verify that the returned trace ID exists in the tracing backend and that the user-facing trace lookup resolves it.
- If using remote write, repeat the query against the destination and compare exemplar visibility.
- Repeat with the feature disabled or remote forwarding disabled in a test environment to confirm which layer is responsible for each result.
Do not validate only by inspecting one successful trace link. Test a normal request, a high-latency request, no-trace sampling, buffer pressure, remote receiver outage, restart, and a version upgrade. This reveals whether exemplars are sparse by design, being dropped at one boundary, or rolling out of the buffer sooner than expected.
Keep correlation useful and bounded
An exemplar should answer a specific operational question: “Which trace can help explain this metric observation?” Keep its labels small and stable. A trace ID is usually sufficient to look up the trace; a span ID can be helpful when the backend supports span-level navigation. Adding duplicate service, tenant, route, and environment labels to every exemplar wastes space if the series already provides those dimensions.
Never attach credentials, session tokens, email addresses, raw query strings, or other sensitive payloads as exemplar labels. Exemplar labels can be exposed through APIs, remote storage, and dashboards even when the trace backend has stricter access controls. Keep the identifiers opaque, access-control the relevant data paths, and ensure the trace backend applies its own retention and authorization policy.
Sampling and retention must agree. An exemplar whose trace was discarded by tail sampling points nowhere; a retained trace without any metric exemplar may still be discoverable through the trace backend’s own search. For critical workflows, document which sampling decision is authoritative, whether the exemplar is emitted only for retained traces, and what operators do when a trace has expired.
Do not use exemplar counts as an application request count or as an unbiased latency sample. The Prometheus sample is the measurement; an exemplar is an optional correlation reference. Querying a narrow time range and checking the underlying histogram or counter remains essential when measuring rates, percentiles, or SLO compliance.
Upgrade and rollback plan
Because exemplar storage and the query endpoint are experimental in current upstream documentation, make the feature an explicit opt-in with a tested off switch. Before upgrading Prometheus, check the version-matched release notes and feature-flag documentation for changes to storage, scrape negotiation, API behavior, and remote-write encoding. A configuration accepted by one release is not a compatibility guarantee for every receiver or query client.
Roll out to one canary Prometheus first. Record exemplar rate, memory use, WAL growth, scrape behavior, remote-write queue health, and trace resolution success. Compare those signals against the same server before enablement. If memory, storage, or receiver behavior is unacceptable, disable exemplar ingestion and forwarding through the planned rollback rather than increasing the buffer without a measured capacity model.
Production checklist
- The deployed Prometheus version explicitly supports the feature flag, exemplar config, exposition format, and query path being used.
- The instrumented application emits exemplars with bounded, non-sensitive trace identifiers, and the scrape actually negotiates a compatible format.
- The buffer limit is sized from measured exemplar rate and memory, with the understanding that it is a rolling capacity rather than fixed-duration retention.
- Remote write is configured separately and enabled only after the receiver’s exemplar support and query behavior are tested.
- A known exemplar can be queried locally, found remotely where required, and resolved to an available trace.
- Sampling, trace retention, exemplar retention, dashboards, and the fallback investigation workflow are documented together.
- Feature-flag disablement, server restart, buffer pressure, receiver outage, and version upgrade have been exercised in a non-production environment.
Exemplars are a small but powerful bridge between aggregate telemetry and request-level traces. They work well when their experimental status is respected, their identity stays out of metric labels, and every hop from instrumentation through storage to trace lookup is verified independently.
Related:
- Prometheus Native Histograms: Sparse Buckets, Mergeable Resolution, and Migration
- OpenTelemetry Tail Sampling: Making Trace Decisions After the Outcome Is Known
Sources: