Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

GKE Workload Identity Federation: Principal Scoping and Metadata Paths

Design GKE workload federation with Kubernetes principals, scoped IAM bindings, metadata-server verification, and a safe migration from service-account keys.

Workload Identity Federation for GKE lets a Kubernetes workload authenticate to Google Cloud APIs without storing a long-lived service-account key in a Secret, image, node disk, or CI artifact. The GKE metadata server mediates token requests for Pods, while Google Cloud IAM evaluates a Kubernetes workload principal or an explicitly linked IAM service account. The feature removes a dangerous credential-distribution pattern, but correct access still depends on node-pool configuration, service-account selection, IAM policy scope, supported workload behavior, and network reachability.

The phrase “workload identity” can refer to several related but different models. GKE supports direct access by Kubernetes service-account principals as well as linking a Kubernetes service account to an IAM service account in some scenarios. These models have different IAM principal strings and operational tradeoffs. Decide which model you are using before writing policies; confusing a Kubernetes service-account principal with an IAM service-account email is a common source of both access denial and excessive permissions.

Understand the metadata path

When federation is enabled for a GKE node pool, the GKE metadata server intercepts supported metadata credential requests made by Pods on that pool. It provides credentials for the Kubernetes service account rather than exposing the underlying Compute Engine instance identity through the ordinary metadata path. This interception is important: if a workload can still obtain a powerful node service account, per-Pod IAM scoping is defeated.

Host-network Pods are a notable boundary. They do not use GKE Workload Identity Federation in the same way as ordinary Pods and can reach a different metadata path. Treat hostNetwork: true as a security-sensitive exception that needs an explicit review rather than an incidental Pod setting. Also verify node-pool coverage: a cluster may contain pools with different federation configuration, and a rescheduled workload can behave differently if it lands on a pool that was not prepared.

Enabling federation does not itself grant access to any Google Cloud API. IAM policy bindings are still required, and their member strings must refer to the intended Kubernetes identity or linked service account. If a principal identifier includes a project number, pool name, namespace, service-account name, or cluster identifier, copy it from the official documented format and verify every component. Avoid hand-building a policy from memory.

Choose a principal model

For direct federation, IAM policies can bind roles to a Kubernetes service-account principal in the GKE workload identity pool. That keeps the Kubernetes namespace and service account visible in the IAM member string and avoids maintaining a separate IAM service account for every workload. The binding should be applied at the narrowest resource that supports the needed role, such as one storage bucket or one secret, rather than at project scope by default.

An alternative is to link a Kubernetes service account to an IAM service account using the documented service-account impersonation relationship. This can ease compatibility with APIs and integrations that expect an IAM service account identity, but it creates another identity object and binding to maintain. The IAM service account’s permissions remain the effective cloud permissions; Kubernetes-level selection of the linked account must be protected by namespace authorization and control of the service-account annotation.

Google documents principal forms for a single Kubernetes service account, all service accounts in a namespace, a service-account UID, or a set of Pods in a cluster. Broader principalSet forms can be useful for shared platform components, but their membership changes as the namespace or cluster changes. UID-based selectors can make identity follow an object instance rather than its reusable name; that can be desirable when deletion and recreation must not inherit the former identity. Choose the narrowest semantics that match the lifecycle and review how namespace reuse or cluster recreation affects the binding.

Configure and verify a service account

The Kubernetes service account is the identity selector for the Pod. Specify it explicitly in the Pod template. For direct federation, the Cloud IAM policy must authorize the corresponding principal. For a linked IAM service account, the Kubernetes service account needs the documented annotation and the IAM service account’s policy must allow the appropriate Kubernetes principal to impersonate it.

apiVersion: v1
kind: ServiceAccount
metadata:
  name: invoice-reader
  namespace: billing
  annotations:
    iam.gke.io/gcp-service-account: "invoice-reader@PROJECT_ID.iam.gserviceaccount.com"
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: invoice-exporter
  namespace: billing
spec:
  replicas: 2
  selector:
    matchLabels:
      app: invoice-exporter
  template:
    metadata:
      labels:
        app: invoice-exporter
    spec:
      serviceAccountName: invoice-reader
      containers:
        - name: exporter
          image: example.invalid/finance/exporter@sha256:REPLACE_WITH_APPROVED_DIGEST
          command: ["/app/export"]

This manifest illustrates the linked-IAM-service-account model and uses placeholders; it is not sufficient by itself. Configure the node pool, grant the documented workload principal permission to impersonate the IAM service account, and grant the IAM service account only the resource permissions required. If using direct federation, omit the annotation and bind the documented Kubernetes principal directly instead. Do not grant both models broadly just to make authentication “work.”

The Pod’s Google client libraries normally use Application Default Credentials, which in GKE can obtain credentials through the metadata server when the node pool is configured correctly. Test with the same image and Kubernetes service account as production. A local developer credential file can mask a broken metadata path, so explicitly verify that the test environment does not mount a service-account key or inherit GOOGLE_APPLICATION_CREDENTIALS pointing at a file.

Scope authorization and prevent node identity escape

Create IAM bindings at resource scope where supported. For a workload that reads one secret, grant a secret-level accessor role on that secret rather than project-wide access to all secrets. For a bucket, scope access to the intended bucket and choose object permissions with the application’s read/write behavior in mind. Workload identity answers “who is calling?”; service-specific IAM and resource policies answer “what may it do?”

The node identity remains a separate principal. GKE node service accounts should receive only permissions required by node components. Do not rely on federation while leaving a broad node identity reachable from Pods through metadata. Validate that the metadata server is running and that tests from ordinary Pods return the workload identity, not the node’s service account. Review hostNetwork, privileged workloads, and node-local access as explicit exceptions in policy.

Namespace names can be reused. If IAM bindings are keyed by a namespace/service-account name, deleting and recreating that object with the same names can recreate a principal with the same textual identity. Restrict who can create service accounts and workloads in identity-bearing namespaces, enforce ownership, and consider UID-scoped principals when object-instance semantics are needed. A GitOps repository that grants a service account access is part of the IAM trust root and deserves code review protections.

Migrate from key files incrementally

Inventory service-account key files, Kubernetes Secrets, projected volumes, environment variables, workload manifests, and application credential configuration. Determine which principal each application currently uses and what APIs it calls. Create the federated IAM binding while leaving the old key path available only for a brief, controlled canary. Remove the key mount and credential-file environment variable from that canary, verify its actual identity and least-privilege calls, then roll the migration through the service.

After all consumers are migrated, disable or delete obsolete service-account keys through the approved rotation process and inspect audit logs for any remaining use. Do not delete a key simply because the new Pod starts: startup may not exercise the cloud API, and a hidden fallback can still use the old file. Run a test that performs the real API call, observe audit logs, and verify the caller principal. Keep the rollback bounded; a long-lived key left in a repository “just in case” undermines the change.

Troubleshoot by layer

If the client cannot obtain credentials, confirm that the Pod is on a correctly configured node pool, is not using host networking, and reaches the GKE metadata server. Verify that node-level federation is enabled and the relevant metadata server component is healthy. Confirm that the application uses supported Google client libraries and has no stale key-file override. Test from inside the workload network and avoid dumping credential responses.

If token acquisition succeeds but an API returns permission denied, inspect the principal in Cloud Audit Logs and compare it with the IAM binding member. Check whether the policy was bound at the right resource, whether the service account annotation points to the intended account, and whether the impersonation relationship is present when linking accounts. A successful token does not mean that a role binding has propagated or that the requested API’s own authorization requirements are met.

When enabling federation across a cluster, use a node-pool rollout plan. Update a small pool or a representative workload first, watch token and API error rates, and retain capacity on the old pool until confidence is established. Reconcile IAM policy changes in a controlled deployment rather than editing production bindings manually during an outage. Document which identity model each workload uses, the target resource scope, the owning team, and how access is revoked.

Federation is a stronger default than distributing service-account keys, but it does not replace workload isolation. Its security value depends on precise principal selection, constrained node credentials, explicit Pod service accounts, and resource-scoped IAM. Validate those properties on the live workload path, not only in a Terraform plan or manifest diff.

Related:

Sources:

Comments