Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

Kubernetes API Encryption at Rest: KMS, Rotation, and Data Rewrites

Encrypt Kubernetes API objects in etcd with an ordered provider configuration, external key protection, measured migration, and a tested decryption recovery path.

Kubernetes API encryption at rest protects selected API objects stored by the API server in etcd. It is separate from encrypting a disk, a cloud volume, or a file mounted into a container. When a Secret is encrypted at rest, the API server encrypts its persisted representation before writing it to storage and decrypts it when authorized clients read it. The encryption provider configuration, key lifecycle, and API server availability therefore become part of the control plane’s data recovery design.

Encryption is not authorization. A user with permission to read a Kubernetes Secret through the API can receive its plaintext from the API server even when etcd stores ciphertext. RBAC, admission policy, workload identity, audit logging, and careful secret distribution remain necessary. Conversely, someone with direct access to raw etcd snapshots may be able to read unencrypted resources unless encryption is correctly configured and existing objects have been rewritten.

Understand the provider chain

An EncryptionConfiguration maps resource types to an ordered list of providers. The first provider is used to encrypt new writes; providers later in the list may be kept to decrypt older data during migration. The identity provider stores data as-is and provides no confidentiality. If it is first, new writes are not encrypted even if an encryption provider appears later in the list.

The available providers have different key ownership and operational properties. Local AES providers store key material on control-plane hosts and require secure distribution and backup. The KMS provider uses envelope encryption with a data encryption key protected by a key-encryption key managed by the external KMS plugin. Kubernetes documents KMS v2 as stable starting with v1.29 and recommends evaluating it for supported clusters. Provider availability differs by control-plane implementation, so managed-service operators must use their provider’s supported feature set.

KMS reduces direct exposure of the data encryption key but adds an external availability dependency. The API server must reach the plugin or remote KMS path to decrypt data. If the key is disabled, deleted, inaccessible, or its plugin is unhealthy, reads of encrypted objects can fail. Keep the key policy, KMS endpoint, plugin health, and key backup or recovery procedure within the control-plane disaster-recovery plan.

Configure encryption before writing sensitive data

For self-managed clusters, the API server needs a readable encryption configuration file and a secure way to access the selected provider. The provider list should put the active encryptor first and retain old decryptors during a planned rotation. Access to the config file and key material should be limited to control-plane administrators. Managed clusters may expose encryption settings as a provider API or may manage them internally; do not assume direct access to an API server host is available.

apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
  - resources:
      - secrets
    providers:
      - kms:
          apiVersion: v2
          name: kms-provider
          endpoint: unix:///var/run/kmsplugin/socket.sock
      - identity: {}

This illustrates the ordered provider shape for a KMS v2 plugin and is not a complete deployment. The socket path, plugin implementation, authentication, timeouts, key policy, and API server flags must match the cluster’s supported KMS integration. Keep identity last as a decrypt fallback for existing plaintext objects during migration; remove it only after the encryption state has been verified. Do not deploy a sample socket path without installing and monitoring a compatible provider.

The resource selector determines what data is encrypted. Restricting the configuration to secrets is common because those objects can contain credentials, but custom resources may also carry sensitive data. If you add resource types or wildcards, evaluate API server cost, provider compatibility, and the impact on startup or control-plane availability. Data already stored in etcd is not automatically transformed simply because the configuration now encrypts writes.

Rewrite existing objects after enabling encryption

After enabling the provider, new writes use the first configured provider. Existing objects remain in their old storage format until they are rewritten. A secret that has never changed since encryption was enabled can therefore remain plaintext in etcd. Run a controlled rewrite for existing resources using the supported Kubernetes procedure, and monitor API server errors, etcd load, KMS request rates, and application impact.

The Kubernetes documentation demonstrates a privileged rewrite operation for existing Secrets:

kubectl get secrets --all-namespaces -o json | kubectl replace -f -

The command reads Secret objects through the API and submits them again so the API server writes them using the active provider. Treat the pipe as sensitive data in transit on the operator workstation: use an isolated administrative host, avoid shell tracing and terminal capture, and ensure command output is not persisted to an unprotected file. For a large cluster, plan load and request sizing, and use the official procedure for the actual control-plane version. A successful API call should be followed by encryption-state verification, not assumed to have covered every object.

Kubernetes exposes storage metrics and provider-specific evidence that can help confirm which resources have been encrypted. Verify with the method supported by your control-plane version, and test reading the same resources through the API after the rewrite. Do not inspect raw etcd data casually on a live cluster; snapshots and database files contain sensitive information and need approved handling.

Rotate keys without losing decryptability

Key rotation is a multi-phase migration. First distribute the new decryption key or KMS key access while retaining the old decryptor. Restart or roll API server processes in the supported sequence so every server can decrypt objects written with either key. Then move the new key or provider to the first position so new writes use it. Rewrite existing objects to migrate their stored data, verify that no object still depends on the old key, and only then remove the old decryptor.

For a highly available control plane, rolling configuration changes must preserve a common decrypt set during every step. If one API server cannot decrypt an object written by another, reads can fail intermittently depending on which server handles the request. Back up the new local key or preserve the KMS key and policy in the recovery account before changing provider order. Do not delete an old KMS key just because new writes use the new key; historical ciphertext can continue to require the old key until all objects have been rewritten.

Run a rotation rehearsal with non-production objects and document the exact config, rollout order, health gates, rewrite procedure, and rollback step. Keep an immutable copy of the last known-good configuration in the control-plane recovery system. Alert when the KMS plugin reports health or latency problems and when API server encryption metrics show an unexpected increase in plaintext or old-key usage.

Preserve backups and restore capability

An etcd snapshot of encrypted API data is useful only if the restore environment can decrypt it. Include the encryption configuration, local keys or KMS access path, plugin version, API server version, and necessary trust material in the recovery plan. Store them separately from the etcd snapshot but within a protected administrative boundary. Protect backups from deletion and test a restore in an isolated cluster.

If a KMS key is unavailable, do not reconfigure the API server to identity and assume encrypted objects will become readable. The data still requires the original decryptor. Recover access to the KMS key or restore the correct key and plugin path first. If local key material is lost, the corresponding ciphertext cannot be decrypted; workloads depending on those objects may fail, and the only safe recovery may be from a backup that includes the key.

Disk encryption and etcd TLS protect different layers. Use transport encryption between API servers and etcd, protect etcd credentials, restrict host access, and encrypt storage volumes as defense in depth. None substitutes for object-level API encryption when a storage snapshot is in scope.

Diagnose encryption failures

If API reads fail after rollout, inspect API server logs and provider health before changing the provider list. Determine whether the error is due to a missing key, KMS authentication, plugin socket, version mismatch, or object format. Confirm that all API server instances have the same ordered configuration and that control-plane restarts completed. Avoid removing old keys until the decryption path is verified for all resources.

If Secret objects appear encrypted in one test but plaintext in a storage inspection, check provider order, resource selector, API server process configuration, and whether the object was rewritten after enablement. If writes fail, inspect KMS latency, quotas, network routes, and API server encryption metrics. Do not make broad changes to all API objects during an incident without establishing the blast radius and restore path.

Kubernetes API encryption at rest is a control-plane cryptographic contract. It requires the correct provider order, complete migration of existing objects, protected keys, API server consistency, and recoverable etcd backups. Validate both confidentiality and decryptability; a system that encrypts everything but cannot restore its keys is not a successful security design.

Related:

Sources:

Comments