Skip to content
SRE & DevOpsDeep Dive Published Updated 10 min readViews unavailable

Kubernetes Image Credential Providers: Exec Plugins, Matching, Caching, and Trust Boundaries

Operate kubelet registry credential plugins with precise image matching, safe caching, workload identity, node rollout controls, and verifiable pull diagnostics.

Kubernetes can ask a node-local executable to obtain registry credentials when the kubelet pulls a container image. The kubelet credential-provider mechanism is an exec-plugin contract: the kubelet selects configured providers whose image patterns match, sends a versioned request over standard input, and reads a versioned response from standard output. A plugin can exchange a node or workload identity for short-lived registry credentials instead of requiring a long-lived registry password to be stored in each namespace.

This mechanism has a different scope and trust boundary from a Pod’s imagePullSecrets. It is configured at kubelet level, applies on nodes where the provider is installed and configured, and can be especially useful for static Pods that cannot refer to namespaced Secrets. It also introduces privileged node software and configuration that must be deployed consistently. Use it only with an implementation maintained by the registry or cloud provider, or with a plugin your organization owns, reviews, signs, and supports.

Decide which credential path fits the workload

Kubernetes supports imagePullSecrets for registry credentials attached to Pods; the referenced Secret must exist in the Pod’s namespace. This is often the simpler mechanism for ordinary workloads when the credential lifecycle and access boundary are namespace-oriented. A kubelet exec provider is a node-level choice when credentials need to be fetched dynamically, have short lifetimes, or should not be persisted as registry credentials in Secrets or on disk.

Avoid treating these paths as interchangeable. A workload Secret is visible to the API and authorized namespace readers according to RBAC; a kubelet provider’s executable and config live on each node and run in the kubelet’s execution context. A provider may grant access to more image namespaces than one Pod requires if its matcher or cloud identity is too broad. Conversely, a node-level provider cannot necessarily express per-workload authorization unless the selected credential design uses a pod identity supported by the cluster and plugin.

Starting with Kubernetes v1.26, the current documentation marks the exec credential-provider mechanism stable. The kubelet invokes a provider only when an image matches its configured patterns. For Kubernetes v1.33 and later, an opt-in service-account-token integration can pass a token bound to the Pod to the plugin when the required feature gate and tokenAttributes configuration are enabled. The documentation describes that token feature as beta since v1.34 and enabled by default there; verify the exact release, feature-gate state, API-server authorization, and managed-cluster support before relying on it.

Install and configure as node software

The provider is an executable binary on every relevant node. The kubelet configuration points to the directory containing binaries and to a CredentialProviderConfig file. A provider’s name must match its executable name, its API version must be supported by the kubelet, and its output must use the same versioned credential-provider API as the input request.

The configuration below is illustrative. The executable name and matchImages path are placeholders for a provider you have built or obtained from a trusted implementation. It does not contain a username, password, or token:

apiVersion: kubelet.config.k8s.io/v1
kind: CredentialProviderConfig
providers:
  - name: registry-example-com-credential-provider
    matchImages:
      - registry.example.com/platform
    defaultCacheDuration: "5m"
    apiVersion: credentialprovider.kubelet.k8s.io/v1

Configure the kubelet’s --image-credential-provider-config and --image-credential-provider-bin-dir using the node platform’s supported method. The executable must be present on each node before a workload requiring it is scheduled there. Package and distribute it through the managed node image, bootstrap system, or provider-supported mechanism; do not copy binaries ad hoc to only the currently running nodes.

A configuration entry also supports arguments and environment variables. The plugin receives the host process environment plus configured variables. Treat both the executable and its environment as part of the node’s privileged supply chain: pin and verify binary versions, review who can modify the directory and config, limit environment values to what is needed, and never put raw secrets in command-line arguments or broadly readable node configuration. Use the plugin’s documented identity flow to acquire credentials at runtime.

Scope image matching precisely

matchImages is an allowlist for when a provider may be invoked, not a registry authorization policy by itself. Patterns can include a registry domain and optional path or port. Globs are supported in domain labels but not in the port or path; a glob matches only one domain segment. The registry domain’s part count must match, the configured path must be a prefix of the target image path, and any configured port must match.

This means *.registry.example.com does not mean every depth of subdomain, and a pattern for registry.example.com/team is narrower than the entire registry. Do not assume that registry.example.com/team and registry.example.com/team-archive represent different security boundaries simply because a human reads them as separate words: the configured path is a prefix matcher, so test the precise image paths the kubelet will request. Conversely, a suffix wildcard that appears broad may match only a single domain label.

Multiple providers can match one image. The kubelet can invoke them and combine credentials; when keys overlap, the earlier provider’s matching value is tried first according to the documentation. Keep provider order intentional and avoid overlapping matchers unless there is a tested reason. A provider that is invoked for a wider namespace than expected may increase node identity usage, registry load, and the chance of returning a credential to an unintended pull.

Before production, test positive and negative image names: exact registry host, a valid subdomain, a nested subdomain that should not match, a path-prefix sibling, and an explicit port. Use a staging node and a registry where the test identity can read only a harmless test repository. Do not infer matcher behavior from a successful image pull alone; that pull could have succeeded through a different provider, a cached credential, an imagePullSecret, or an image already on disk.

Align credential lifetime with cache lifetime

Each provider declares defaultCacheDuration, the in-memory credential cache lifetime used when its response does not provide a duration. This is required configuration, not a recommended universal value. The plugin can return a cache duration appropriate to its credentials. If the credential expires in five minutes, allowing the kubelet to cache it for hours can turn an initially successful pull into later authentication failures. If credentials are fetched on every image pull without a reason, the registry identity service can become a new availability dependency and a burst of node starts can amplify calls.

Choose the cache duration from the returned credential’s actual expiry and revocation behavior. Leave room for clock skew, pull retries, and the time required to complete a large image download. Test a pull started near expiry and a later pull after the provider must refresh credentials. Document whether revocation takes effect immediately, only after cache expiry, or through another provider-defined mechanism; the kubelet cache is not a revocation service.

When the plugin uses a Pod-bound service-account token, the cache key changes the isolation boundary. The Kubernetes documentation describes Token as the conservative cache type for credentials whose lifetime is limited to that token. ServiceAccount can be appropriate when the credentials are valid for all Pods using the same service account and plugin authorization depends on the service account rather than Pod-specific claims. Do not select ServiceAccount merely to improve cache hit rate if the registry credential actually carries per-Pod authorization. The API reference notes that metadata and configured service-account annotation keys affect cache dimensions; understand exactly what differentiates entries before choosing a shared cache scope.

For pass-through designs that must not share a credential across Pods, Kubernetes documents returning a zero cache duration as a way to disable credential caching; this causes the plugin to run for each pull and adds token generation and provider-call overhead. Test both authorization isolation and the extra control-plane/identity-service load before adopting that pattern.

Workload identity is not automatic

The optional service-account-token path can let the plugin exchange a Pod-bound token for registry credentials. It does not mean that the registry understands Kubernetes tokens directly: the plugin and an identity exchange must validate the intended audience, issuer, subject and any claims they rely on. Keep the token audience narrow, use TLS-protected identity endpoints, and follow the registry or cloud provider’s documented trust configuration.

The tokenAttributes configuration includes the intended audience, a cache key type, and whether a service account is required. If the plugin requires specific ServiceAccount annotations, the configured required keys must exist or the kubelet will not invoke the provider and returns an error. Optional annotation keys still require the plugin to validate their values. Do not let users set an annotation that silently selects a more privileged registry identity without admission controls and a clear authorization model.

The feature gate must be enabled when tokenAttributes is configured; Kubernetes documents that kubelet fails to start with invalid configuration if those fields are set while the gate is disabled. On clusters where service-account audience authorization restrictions are enabled, the node identity may also need permission to request tokens for the configured audience. A successful YAML parse is therefore not proof that the kubelet can obtain a token or that the exchange is authorized.

Diagnose image-pull failures by layer

Start by collecting the exact image reference, Pod events, node assignment, and image-pull error. Then verify the node has the expected provider binary and configuration, the plugin name matches the executable, the kubelet flags point to the intended files, and the image matches the provider’s configured pattern. Inspect kubelet and provider logs through the operating system or managed-node support path; avoid printing credentials or token contents into shared logs.

Separate matching and execution errors from registry authorization failures:

  1. If the provider is not invoked, check the registry host, path, port, wildcard depth, and whether the image is already cached.
  2. If it is invoked but exits unsuccessfully, validate binary architecture, executable permissions, dynamic-library requirements, stdin parsing, stdout JSON, and stderr diagnostics.
  3. If the plugin returns credentials but the registry denies the pull, verify the credential’s audience, scope, repository permissions, expiry, and clock assumptions.
  4. If the first pull succeeds but later pulls fail, check provider response duration, defaultCacheDuration, credential expiry, node time, and whether cached credentials are being reused past their useful lifetime.
  5. If only some node pools fail, compare kubelet versions, provider binary checksums, config distribution, node identity, and runtime-specific support instead of changing Pod YAML at random.

Use a dedicated test image and a controlled identity to verify an unauthenticated negative case, a permitted pull, a forbidden repository, a refresh after expiry, and a node replacement. A rollout test should include newly created nodes because a warm node can hide a missing binary or broken identity path until autoscaling or disaster recovery occurs.

Roll out and operate with a clear boundary

Treat changes to provider binaries, matching rules, identity policy, and cache duration as security-sensitive node-pool changes. Version the config with the node image or its source of truth, canary a small pool, confirm a fresh node can pull a private test image, and define an immediate rollback to the prior known-good provider. Preserve event and log evidence for failures, but redact bearer tokens, password fields, and authorization headers.

Monitor provider invocation errors, credential exchange latency and error rates, registry authentication failures, image-pull duration, and the rate of new node starts. Avoid high-cardinality labels such as raw image digests or Pod UIDs in generic metrics unless the observability system is deliberately sized for them. Alert on user-impacting pull failures and on identity-provider outages that can prevent new workloads from starting; existing running containers may continue while new Pods cannot be created.

Use the built-in provider when a workload’s namespace-scoped imagePullSecrets already meet the requirement. Use an exec provider when node-level integration or dynamic short-lived credentials justify a managed executable on every node. Add Pod-bound tokens only when the identity exchange and cache isolation are explicitly designed and supported. In all cases, prove the path on a cold node and verify that the plugin grants no more registry access than the intended workloads need.

Related:

Sources:

Comments