OpenTelemetry k8sattributes: Reliable Kubernetes Metadata Enrichment
Enrich OpenTelemetry telemetry with Kubernetes metadata using stable pod identity, explicit association rules, narrow RBAC, and bounded collector caches.
The OpenTelemetry Collector Kubernetes Attributes Processor, commonly configured as k8sattributes, enriches spans, metrics, and logs with metadata read from Kubernetes objects. It first has to associate an incoming telemetry resource with the Pod that produced it. If that association is wrong or missing, a beautifully formatted set of labels can still misidentify the workload, silently omit context, or increase backend cardinality without improving diagnosis.
Production enrichment therefore starts with identity and topology, not with a long list of labels. Decide how telemetry will identify its source Pod at each hop, configure explicit association rules, grant the collector only the Kubernetes API access it needs, and extract a small set of useful attributes. Then test both successful matches and the paths that intentionally remain unmatched.
Separate association from extraction
The processor performs two different jobs:
- Association matches telemetry to a Kubernetes Pod using a connection property or resource attributes already present on the telemetry.
- Extraction reads selected metadata from that Pod and related Kubernetes objects, then adds the requested values as OpenTelemetry resource attributes.
Association does not infer workload ownership from a service name, a metric name, or a Prometheus label that happens to look familiar. It needs an input the processor can compare with its Pod cache. The default behavior can use the incoming connection’s source IP, but that is reliable only when the collector receives traffic directly from the workload through a network path that preserves the useful Pod address.
For an in-cluster gateway, prefer a stable resource attribute when the instrumentation or an upstream agent can provide one. A Pod UID is less vulnerable than a Pod name or IP to reuse over time. OpenTelemetry’s Kubernetes semantic conventions identify k8s.pod.uid as the Pod identity attribute; names are easier for humans to read but can be reused after an object is deleted. An IP can be useful, but NAT, proxies, sidecars, and address reuse can make it ambiguous.
A compact configuration can associate first by Pod UID and fall back to an available Pod IP, then extract a bounded set of metadata:
processors:
k8sattributes:
auth_type: serviceAccount
pod_association:
- sources:
- from: resource_attribute
name: k8s.pod.uid
- sources:
- from: resource_attribute
name: k8s.pod.ip
extract:
metadata:
- k8s.namespace.name
- k8s.pod.name
- k8s.pod.uid
- k8s.deployment.name
labels:
- from: pod
key: app.kubernetes.io/name
tag_name: k8s.pod.label.app.kubernetes.io/name
This snippet assumes the incoming telemetry already has one of the configured identity attributes. The two source lists are ordered fallback rules: the processor tries the UID association first and uses the IP rule if the earlier rule does not match. Within a single association rule containing multiple sources, all of those sources must match. Validate these details against the Collector Contrib version actually deployed, because processor configuration evolves with releases.
The metadata list is deliberately short. Extracting Pod name, namespace, UID, and Deployment name usually establishes useful workload context, but the right set depends on the questions operators need to answer. Label extraction is opt-in and should name only approved keys. Avoid copying every Pod label or annotation into telemetry by default; those values can contain sensitive or high-cardinality data and may be indexed as dimensions by downstream backends.
Place the processor where identity still exists
The connection-based association rule depends on connection context. Put the processor before components that discard that context, such as batching or tail sampling, when the processor uses the connection source. A configuration that works when the receiver connects directly to a Pod can fail after an agent, service mesh, proxy, or gateway changes the observed peer address.
In an agent-and-gateway topology, the gateway normally sees the agent as the network peer, not the originating application Pod. Do not ask the gateway to infer the original Pod from the agent’s IP. There are two common designs:
- Run full Kubernetes association and extraction in a node-local agent, where the agent observes workload traffic and can be filtered to Pods on its own node.
- Have the agent preserve a trustworthy Pod identity attribute, then configure the gateway to associate using that resource attribute while it performs centralized enrichment.
The Collector Contrib documentation also describes a passthrough mode for agents that need to forward identity without querying Kubernetes or extracting metadata locally. In that design, confirm which identity attribute is produced and preserved, and verify that the gateway can resolve it against its own cache. A processor cannot add metadata based on an attribute that was dropped upstream.
Pod association rules are evaluated in order and stop when a rule matches. Use specific, unambiguous identity sources before weaker fallbacks. If the first rule matches the wrong Pod, a later rule cannot correct the result. Emit diagnostics for unmatched telemetry and sample them during rollout; an empty attribute is not the same as an intentional “unknown workload” value.
Use least-privilege Kubernetes API access
The processor watches Kubernetes objects to maintain an in-memory metadata lookup. Its RBAC requirements follow the data it discovers and extracts. A cluster-wide setup that looks up Pods across namespaces generally needs get, list, and watch on the relevant Pods and namespaces. Extracting Deployment ownership or labels may require ReplicaSet permissions; requesting Node or Node label metadata can require access to Nodes. Add permissions only for enabled features and verify the actual API calls against the Collector version and configuration.
For a collector limited to a known workload namespace, a namespace filter can reduce both the lookup scope and the permissions needed:
processors:
k8sattributes:
auth_type: serviceAccount
filter:
namespace: payments
pod_association:
- sources:
- from: resource_attribute
name: k8s.pod.uid
extract:
metadata:
- k8s.namespace.name
- k8s.pod.name
- k8s.pod.uid
Bind a dedicated collector ServiceAccount to a namespace Role rather than granting broad access by habit:
apiVersion: v1
kind: ServiceAccount
metadata:
name: otel-collector
namespace: observability
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: otel-pod-metadata-reader
namespace: payments
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["replicasets"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: otel-pod-metadata-reader
namespace: payments
subjects:
- kind: ServiceAccount
name: otel-collector
namespace: observability
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: otel-pod-metadata-reader
This example assumes the processor needs Pod and ReplicaSet data in payments and does not need cluster-scoped Node or Namespace metadata. Add other resources only when the configured extraction requires them. With namespace-scoped RBAC, the processor cannot query cluster-scoped objects such as Nodes and Namespace labels. The current Contrib documentation also notes that k8s.cluster.uid is derived from the UID of the kube-system Namespace, so a namespaced Role cannot read it unless additional cluster-scoped access is granted.
For a DaemonSet agent that discovers local workloads, set a node filter from the Pod’s spec.nodeName using the Downward API and configure filter.node_from_env_var. Without that filter, every agent may watch far more Pods than it needs. This can increase API traffic and cache memory with little operational benefit. Test the ServiceAccount and filter together; a role that is too narrow or a node name that is unset can look like an association failure even when telemetry contains a valid identity.
Keep cache size and enrichment cost bounded
The processor caches metadata for Pods it watches. The memory cost grows with the discovered workload set and the amount of metadata retained. A gateway that watches every Pod in a very large cluster can have materially different resource needs from a node-local agent filtered to one machine. Measure process memory and Kubernetes API behavior under realistic Pod churn before scaling the Collector based only on incoming telemetry volume.
Use namespace and node filters where they match the deployment topology. Extract only attributes used in dashboards, alerts, routing, or incident investigation. Avoid labels such as per-build IDs, arbitrary release hashes, user-controlled values, or timestamps unless a specific query justifies their dimensional cost. A label that helps identify one Pod may become an expensive metric series dimension when attached to every observation.
Pod deletion and delayed telemetry also require a policy. The processor can retain metadata for a configured grace period after a Pod deletion event so late-arriving data can still be enriched. The current processor documentation lists a default grace period of 120 seconds; verify the exact behavior for the pinned Collector release and set a deliberate value for the application’s telemetry delay and memory budget. A long grace period helps with late data but retains more stale cache entries during high churn.
For very large clusters, review the processor’s informer resync behavior and Collector release notes before changing it. The current Contrib README warns that periodically reprocessing a large informer cache can create CPU and garbage-collection spikes, while Kubernetes watch events already propagate object changes. Treat cache and resync settings as versioned runtime behavior, and load-test changes instead of copying a value from a different cluster size.
Preserve the meaning of resource attributes
Kubernetes metadata enrichment is not a substitute for instrumentation identity. Keep service.name, service.namespace, service.version, and service.instance.id aligned with their OpenTelemetry meanings, and use k8s.* attributes for Kubernetes objects. A service can span multiple Pods, and a Pod can contain multiple containers; mixing those identities makes service-level dashboards and per-Pod investigation difficult to reconcile.
Use the current OpenTelemetry semantic convention names rather than inventing aliases such as pod_name, namespace, or cluster. Consistent keys make it possible to query telemetry from multiple languages and Collector pipelines without a translation rule for each team. If the organization already has legacy keys, normalize them at a documented boundary and avoid emitting two competing values for the same identity.
Only extract labels and annotations that have an owner, a defined value set, and a query use case. Kubernetes labels are operational metadata, not automatically trustworthy telemetry fields. Treat them as input that may change with deployment revisions. Do not use an extracted annotation as a security decision or as the sole source of environment identity unless a separate control validates its provenance.
Diagnose missing or incorrect matches
When enrichment is absent, debug the pipeline in this order:
- Inspect the resource attributes entering the processor. Confirm that the exact key and value used by the association rule exist before enrichment.
- Confirm that the processor watches the Pod and namespace and has the required
get,list, andwatchpermissions. - Check whether network proxies, host networking, an agent, or a gateway changed the connection address.
- Verify the Pod UID, Pod IP, namespace, and object lifecycle at the time the telemetry was produced.
- Check whether the processor was placed after batching or tail sampling when it depends on connection context.
- Confirm that the requested metadata source is enabled and that related object permissions are present for Deployment, ReplicaSet, Job, Node, or Namespace extraction.
- Compare unmatched-rate and cache-size behavior before and after each change.
A configured resource-attribute rule does not mean the processor invents the value. If an SDK, file receiver, or upstream agent never sets k8s.pod.uid, that rule cannot match. Likewise, a successful Kubernetes API watch does not prove every telemetry stream is associated. Keep a sample of raw input and enriched output in a safe test environment so rules can be validated independently of production dashboards.
Production validation checklist
- The association key exists upstream and represents the workload identity intended by the backend.
- UID-based matching is preferred where available; IP and name fallbacks are tested for reuse, proxy, and namespace ambiguity.
- Connection-based association runs before processors that remove connection context.
- Agent-to-gateway pipelines preserve a Pod identity rather than relying on the agent’s peer IP.
- Namespace or node filters align with the topology and bound the Pod cache.
- The ServiceAccount has only the API verbs and resources needed for the enabled metadata features.
- Labels and annotations are explicitly selected, governed, and checked for cardinality and sensitive values.
- Deletion grace and informer behavior are tested under workload churn on the pinned Collector version.
- Unmatched telemetry, memory use, API errors, and processor version are observable.
- Rollout tests cover both correct matches and an intentional unmatched case before relying on enriched dimensions.
The k8sattributes processor is reliable when its association evidence is trustworthy and its metadata scope is deliberate. Start with one stable identity, watch only the Kubernetes objects required to resolve it, and add attributes only when operators can explain how they improve a query. That keeps enriched telemetry useful without turning the Collector into an unnecessarily broad API client or an uncontrolled source of new dimensions.
Related:
- OpenTelemetry Collector Pipelines: Receivers, Processors, Exporters, and Failure Control
- OpenTelemetry Collector Resilience: Memory, Queues, Retries, and Backpressure
Sources: