Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

OpenTelemetry Target Allocator: Prometheus Scrape Ownership at Scale

Scale OpenTelemetry Prometheus scraping with the Target Allocator, explicit monitor selectors, collector identity, RBAC, and duplicate-target diagnostics.

The OpenTelemetry Operator’s Target Allocator (TA) separates Prometheus target discovery from metric scraping. The allocator discovers configured Prometheus targets and distributes them across OpenTelemetry Collector instances. Each Collector receives a target-specific HTTP service-discovery view, allowing a pool of Collectors to share scrape work without every replica scraping every target. The feature is especially useful when a Collector fleet needs to scale horizontally, but it introduces another service, API permissions, selectors, and allocation state that must be monitored.

The Target Allocator is optional. Enabling the targetAllocator section alone does not necessarily enable Prometheus custom-resource discovery; both the allocator and prometheusCR discovery must be configured for ServiceMonitor and PodMonitor objects. The Operator modifies the Prometheus receiver configuration so that it uses HTTP service discovery from the allocator. Existing static or file-based discovery configuration can be replaced as part of that transformation, so inspect the final reconciled Collector config before turning it on.

Separate discovery from scrape execution

Without a Target Allocator, each Prometheus receiver may independently discover and scrape the same targets. That can be intended for redundant independent monitoring, but it multiplies scrape load and duplicate series. The Target Allocator gives each Collector a subset of targets and can redistribute assignments as the collector pool changes. The Collector still owns scrape scheduling and metric processing; the allocator does not store metric samples or export them to the backend.

This split makes scaling more predictable. Adding Collector replicas can distribute target ownership, while adjusting the allocator’s strategy can balance target counts across instances. Equal target count is not necessarily equal work: one endpoint may expose a large, high-cardinality metrics payload while another is tiny. Measure scrape duration, samples, memory, CPU, and exporter queue behavior rather than assuming an even number of targets guarantees even resource use.

Prometheus Custom Resources such as ServiceMonitor and PodMonitor have selectors and namespace selectors. The target allocator must discover only the intended objects, and the Collector’s service account needs permissions to read those resources through the Kubernetes API. A selector that is too narrow silently excludes targets; a selector that is too broad can scrape unintended namespaces or multiply data collection.

Enable the component explicitly

An OpenTelemetryCollector custom resource can enable the allocator and its Prometheus custom-resource discovery. The exact API version and fields depend on the installed OpenTelemetry Operator and Collector CRD. Keep the configuration versioned with the operator release and inspect the resulting Target Allocator Deployment and Service after reconciliation.

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: metrics-collector
  namespace: observability
spec:
  mode: statefulset
  replicas: 3
  targetAllocator:
    enabled: true
    serviceAccount: metrics-target-allocator
    prometheusCR:
      enabled: true
      serviceMonitorSelector:
        matchLabels:
          monitoring: enabled
      podMonitorSelector:
        matchLabels:
          monitoring: enabled
  config:
    receivers:
      prometheus:
        config:
          scrape_configs:
            - job_name: collector-self
              scrape_interval: 30s
              static_configs:
                - targets: ["0.0.0.0:8888"]
    exporters:
      debug:
        verbosity: basic
    service:
      pipelines:
        metrics:
          receivers: [prometheus]
          exporters: [debug]

This is a representative Operator configuration, not a complete production collector pipeline. The custom-resource API version, prometheusCR selector fields, config representation, mode, and service-account name must match the installed Operator CRD. Validate the manifest against that CRD. The debug exporter is shown for configuration clarity; use an approved production exporter and secure its credentials rather than shipping metrics to a debugging sink.

The self-scrape job illustrates a Prometheus receiver configuration that the Operator can transform when the allocator is enabled. The Operator can remove existing service-discovery sources from the configured scrape jobs and add http_sd_configs that point to the allocator. Do not assume a static_configs entry remains active unchanged. Review the rendered Collector configuration and test the exact behavior for your Operator version.

Scope ServiceMonitor and PodMonitor discovery

Use labels and namespace selectors to define the monitor objects the platform intends the collector pool to scrape. Assign monitor ownership through code review and namespace RBAC. A workload team should not be able to add a monitor that causes a privileged telemetry Collector to reach arbitrary internal endpoints unless that is part of the platform contract. Monitor resources can specify endpoints, paths, TLS settings, and relabeling that affect both network exposure and label cardinality.

Decide which team owns target discovery and which team owns the Collector. A central platform may own the Target Allocator and collector service accounts while application teams own their ServiceMonitor objects. Document allowed labels, namespaces, endpoint schemes, and authentication requirements. Use NetworkPolicy and service-mesh controls to constrain Collector egress, while permitting the legitimate scrape paths.

The allocator’s service account needs read access to the relevant monitor APIs and discovery resources. The Collectors need permission to reach their assigned targets and, depending on configuration, query the allocator for target lists. Apply only the RBAC documented for the Operator release and inspect actual API calls. Avoid granting cluster-admin to the allocator to overcome a missing permission; discover which resource and namespace scope is required.

Choose Collector mode and plan rebalancing

The Target Allocator integrates with a Collector custom resource that can run in a supported mode such as Deployment or StatefulSet. Select a mode based on identity, stable network endpoints, storage, and rollout needs. The allocator assigns scrape targets to Collector Pods; when replicas are added, removed, or replaced, target ownership may shift. Scrape intervals and backend deduplication behavior determine whether that shift produces temporary gaps or overlapping samples.

For high availability, run multiple Collector replicas and define how target assignment behaves during a replica failure. Avoid assuming that replicas make a single Target Allocator instance highly available; inspect the Operator-managed topology and its readiness probes. When the allocator is unavailable, existing Collector assignments may continue temporarily or fail to refresh depending on the integration version. Test allocator restart and Collector rollout behavior before relying on it during a monitoring incident.

Use graceful termination and adequate pod disruption settings so a Collector is not removed before it can stop scraping and flush exporter queues. Ensure the Collector’s memory limiter and exporter sending queue are configured for the expected telemetry rate. Rebalancing targets does not guarantee exporter capacity; one collector may still exceed memory or downstream throughput because its targets emit significantly more samples.

Verify allocation and detect duplicate scraping

Start by verifying that the Operator created the Target Allocator Deployment and Service, that the allocator Pod is Ready, and that the Collector has the generated http_sd_configs. Then inspect the allocator’s service-discovery endpoint using the documented diagnostic path and check whether the expected ServiceMonitor and PodMonitor objects appear. If target lists are empty, confirm both feature flags are enabled, selectors match labels, namespace selectors include the monitor namespaces, and RBAC permits reads.

Check target ownership across all Collector Pods. Compare assigned target sets and ensure the same endpoint is not intentionally scraped by both the allocator-managed job and a separate static configuration. Duplicate scraping can result in duplicated series, inflated counters, higher cardinality, and unnecessary network load. Use a stable job or collector identity label where backend queries need to distinguish replicas, but avoid adding pod-unique labels to every metric if that causes unbounded series growth.

If targets are missing, inspect ServiceMonitor endpoint port names, service selectors, endpoint readiness, TLS configuration, and scrape path. The Target Allocator cannot repair an incorrect Kubernetes Service or a network route blocked by policy. If only some metrics disappear, inspect relabeling rules and collector pipeline processors as well as target allocation.

Operate upgrades and failure recovery

Upgrade the OpenTelemetry Operator and Collector versions through a compatibility matrix. Custom-resource fields can change, the Operator may rewrite Collector config differently, and Prometheus receiver behavior can evolve. Before upgrading, render the CRD schema, preserve the current generated collector configuration, and test target count, label set, scrape rate, and exporter queue in a representative cluster. Avoid combining an operator upgrade with a major change to monitor selectors.

Keep an operational fallback. If the allocator is unavailable or its generated configuration is invalid, know how to pause the affected collector change or temporarily restore a known-good configuration. Do not make a blind switch to per-replica static discovery if that would cause every Collector to scrape every target. Capture the current monitor inventory and assignments before changing replica counts.

The Target Allocator is a control-plane component for telemetry collection, not a metrics backend. Its value is clean division of discovery and scrape work, but that boundary adds operational responsibility. Explicit selectors, least-privilege RBAC, generated-config inspection, rebalance testing, and duplicate-target monitoring make horizontally scaled scraping understandable and recoverable.

Related:

Sources:

Comments