Kubernetes API Audit Policy: Coverage, Privacy, and Reliable Delivery
Design Kubernetes API audit policies that preserve useful identity and change evidence while controlling sensitive payloads, event volume, and delivery loss.
Kubernetes API auditing records a chronological trail of requests handled by the kube-apiserver. It can help an operator determine which authenticated identity attempted an API action, which object or URL was targeted, what response status was returned, and when the request passed through the API server. It is not the same thing as a Kubernetes Event, and it is not a host-level system-call log. Treat it as one control-plane evidence source with a deliberately designed policy and a durable, access-controlled destination.
Audit quality is a balance. Recording only a few resource types can leave gaps in an incident timeline; recording request and response bodies indiscriminately can expose credentials and application data, increase API-server memory and storage costs, and make the evidence harder to protect. A production policy should begin with the questions responders must answer, record the least sensitive data that answers them, and be tested under the request volume and failure behavior of the actual control plane.
Define what the API server records
An audit event is generated as an API request moves through the API server. The event can include an audit ID, the authenticated user, impersonated identity when applicable, source IP information, request URI, Kubernetes verb, object reference, response status, timestamps, and audit annotations. Request and response objects are included only at levels that permit those bodies. The userAgent field is supplied by the client and must not be treated as trusted identity evidence. Source IP fields can include forwarding headers, so interpret them in light of the trusted proxy chain rather than as cryptographic proof of origin.
Audit stages describe request progress, not the API operation’s business meaning:
RequestReceivedis emitted when the audit handler first receives the request.ResponseStartedis emitted after response headers are sent for a long-running request, such as a watch.ResponseCompletemarks completion of the response body.Panicrecords a request path that ended in a panic.
One API request can therefore produce more than one stage event. A watch may remain open and generate ResponseStarted before it eventually completes. If RequestReceived is omitted to reduce duplicate records for ordinary requests, do not assume every request now produces exactly one event: long-running requests have different stage behavior. Preserve auditID and stage when indexing or correlating records so separate events for the same request are not mistaken for separate client actions.
The audit API is also distinct from the core events API. A Pod Warning Event is a controller or node report about an observed condition; an audit record is generated from a request received by the API server. Neither source alone is a complete incident record. A direct process action on a node, a syscall inside a container, or an application request that does not pass through the Kubernetes API requires evidence from the relevant host, runtime, or application system.
Choose the audit level deliberately
The policy assigns the first matching rule’s level to a request. Rules are strictly ordered, and every field present in a rule must match. If no rule matches, the default level is None unless a catch-all rule is provided. A later broad rule cannot override an earlier narrow rule, so place intentional exclusions before broader capture rules and review exclusions as carefully as inclusions.
| Level | Evidence recorded | Operational trade-off |
|---|---|---|
None |
No event for a matching request | Lowest volume, but the activity is absent from this audit stream |
Metadata |
Request identity and metadata, without request or response bodies | A practical default for broad coverage with lower data sensitivity |
Request |
Metadata and request body, but not response body | Useful only where the submitted body is necessary evidence and safe to retain |
RequestResponse |
Metadata, request body, and response body | Highest detail and highest sensitivity, size, and storage exposure |
For non-resource requests, the Request and RequestResponse levels do not include request or response object bodies. For resource requests, bodies can contain credentials or other sensitive values. A Secret object, a workload manifest embedding environment values, or a custom resource carrying application data may be captured at Request or RequestResponse. Response bodies can reveal the persisted version of an object, but can also repeat sensitive material. Prefer Metadata for broad coverage and use body capture only for narrowly scoped resources after a data-protection review. omitManagedFields can reduce noisy managed-field data in bodies where supported by the target API server; it does not sanitize other fields or make body logging safe by itself.
Start with a privacy-conscious policy
The following compact policy records metadata for all requests and omits the initial stage. It illustrates explicit broad coverage without recording object bodies. It is a starting point for a test cluster, not a universal production profile: high-volume clusters still need measured capacity, retention, and data-access controls.
apiVersion: audit.k8s.io/v1
kind: Policy
omitStages:
- RequestReceived
omitManagedFields: true
rules:
# The catch-all is intentionally last. It records metadata, not bodies.
- level: Metadata
Before deployment, decide whether the policy must include RequestReceived for specific investigative questions, and whether long-running requests need both later stages. If a narrowly scoped rule requires request-body evidence, insert it before the catch-all, specify the API group, resource, verb, and namespace as tightly as possible, and document the sensitive fields that may be present. Resource names alone do not make a body safe. Re-test rule ordering whenever a policy changes; a broad rule placed too early can make later rules ineffective.
The audit policy is a local kube-apiserver configuration file, not an object submitted to the Kubernetes API. For a self-managed control plane, --audit-policy-file points to that file. In a static-Pod control plane, the policy must be mounted where the API server can read it. Managed Kubernetes services expose audit configuration and delivery through provider-specific controls; do not assume that editing a node or supplying a kube-apiserver flag is supported. Confirm the provider’s current audit-log configuration and retention behavior before relying on it.
Select a backend and failure behavior
The built-in log backend writes JSON Lines to a file. Its path and rotation controls include --audit-log-path, --audit-log-maxage, --audit-log-maxbackup, and --audit-log-maxsize. A file on a control-plane host is not automatically durable: persist it, restrict who can read it, monitor disk capacity, and export it before rotation or node replacement removes the evidence. If the API server runs as a Pod, the file path and policy path must be mounted into that Pod for the intended persistence and configuration to work.
The webhook backend sends audit events to a remote HTTP API configured through a specialized kubeconfig. Its batching mode can buffer events and deliver them asynchronously, while blocking modes couple request handling more directly to audit export. A batch buffer can overflow when incoming events exceed the configured capacity; the API server exposes audit event and audit error counters, including apiserver_audit_event_total and apiserver_audit_error_total, to help observe export behavior. Blocking and blocking-strict modes have different availability and latency consequences. In strict blocking mode, an audit failure at the RequestReceived stage can fail the API request itself. Do not select a mode from a generic performance recommendation: test backend outages, slow responses, queue saturation, and control-plane load in a representative environment.
For either backend, document the failure contract: whether events can be dropped, whether API requests can wait or fail, who receives alerts, and how recovery is verified. Monitor audit export errors and storage saturation alongside API-server latency and availability. A successful API call does not prove that its audit record reached the final retention system; verify the downstream search or archive using known test events.
Protect the audit trail as sensitive operational data
Audit metadata can reveal usernames, service-account identities, namespace and object names, access patterns, request paths, source addresses, and authorization outcomes. Body levels add the full privacy and credential risks of the API objects being submitted or returned. Treat audit storage as sensitive even when the policy uses only Metadata.
Use a dedicated destination with narrowly scoped reader access, encryption in transit and at rest, and retention tied to incident response and applicable records requirements. Keep audit data separate from mutable application logs where practical. Restrict the ability to alter or delete the exported trail, record administrative access to it, and test that timestamps remain usable across control-plane and backend components. Avoid putting bearer tokens or secret values into dashboards, example queries, or incident tickets copied from raw audit records.
Audit records have evidentiary limits. The API server authenticates the request and writes the user field, but a username may represent a shared service account rather than a human. impersonatedUser is important when impersonation is used. A client-supplied user agent is descriptive, not an identity credential. sourceIPs can include X-Forwarded-For and X-Real-Ip values; Kubernetes documents that all but the last IP can be arbitrarily set by the client. Only interpret forwarded addresses in the context of the trusted proxy chain. Preserve the original event and its context rather than turning a single field into a stronger attribution claim than it supports.
Validate coverage before depending on it
Roll out policy changes through the control-plane configuration owner and use the supported procedure for the cluster. In a non-production environment, generate a small known set of requests using a dedicated identity: a successful read, an allowed mutation, a denied request, an impersonated request if that feature is used, and a long-running watch. Confirm that the selected level produces the expected fields and stages, that excluded requests are intentionally absent, and that no request or response body appears under a metadata-only rule.
Then interrupt or slow the backend in a controlled test. Observe whether the API server buffers, blocks, drops records, increments audit errors, or rejects requests under the selected mode. Check disk rotation and retention, downstream ingestion delay, searchability by auditID, and recovery after the backend returns. A policy file passing YAML parsing is not proof that the control plane loaded it or that the chosen backend retains the resulting records.
For an incident, correlate audit records with Kubernetes Events, API object state, controller and node logs, and application telemetry. Use audit data to establish what the API server observed, not to infer a container’s internal behavior or to claim that an object change successfully affected every data plane. This separation keeps the timeline useful and the conclusions proportional to the evidence.
Related:
- Kubernetes RBAC Impersonation: Testing Authorization Without Sharing Credentials
- Kubernetes Events as Operational Evidence: Retention, Aggregation, and Durable Diagnosis
Sources: