Skip to content
SRE & DevOpsHow-To Published Updated 4 min readViews unavailable

Kubernetes Ephemeral Containers: Debugging a Pod Without Rebuilding Its Image

Use kubectl debug and ephemeral containers for live Pod diagnosis while accounting for immutable specs, runtime targeting, access, and cleanup.

An ephemeral container is a temporary debugging container added to an already-created Pod through the Pod’s ephemeralcontainers subresource. It is useful when kubectl exec cannot help because the application container has crashed or its minimal image contains no shell, package manager, or diagnostic utilities. The feature is not a general-purpose way to add sidecars to a live workload: ephemeral containers have deliberately narrower lifecycle and configuration semantics than ordinary containers.

Kubernetes marks ephemeral containers stable from v1.25. They are not automatically restarted, cannot be removed or changed after being added, and do not support fields such as resource requests and limits, ports, or container probes. Plan them as an incident-response tool and keep the Pod disposable or replaceable after the investigation.

Triage before attaching a debugger

Start with evidence that does not mutate the Pod. Capture its state, events, previous container logs, image digest, and node placement:

kubectl describe pod "$POD" -n "$NAMESPACE"
kubectl logs "$POD" -n "$NAMESPACE" -c "$CONTAINER" --previous
kubectl get pod "$POD" -n "$NAMESPACE" -o wide

If the application is still running and already contains the required tools, kubectl exec may be sufficient. If it is repeatedly crashing, an ephemeral container can remain available alongside the existing Pod while you inspect its shared process namespace, if the runtime and Pod configuration support targeting that namespace.

Add an ephemeral container with kubectl debug

Use an approved, pinned debugging image from your organization’s registry. Avoid relying on a floating latest tag during an incident; the same command should produce the same toolset and provenance each time.

kubectl debug -it "$POD" \
  -n "$NAMESPACE" \
  --image=registry.example.net/ops/debug-tools@sha256:REPLACE_WITH_APPROVED_DIGEST \
  --target="$CONTAINER"

--target asks the container runtime to target the process namespace of the named container. Kubernetes documentation notes that this depends on runtime support. If unsupported, the debugger may start with an isolated process namespace, so an empty ps result does not prove the application has no processes. Verify what the cluster version and runtime actually did with kubectl describe pod and the container state.

The operation requires authorization to update the Pod’s ephemeral-container subresource. Access to a debugging image or cluster is not by itself permission to attach to production processes. Use an incident identity, a recorded reason, and the least-privileged debug profile or security context supported by the deployed kubectl and cluster version. Do not default to privileged access simply because it exposes more diagnostics.

Inspect the Pod after the session

Record the exact command, image digest, time, operator identity, and changes made. The ephemeral container is visible in the Pod specification and status, and the container is not a hidden session. Collect output through approved logging or incident tooling; avoid dumping environment variables, mounted secrets, or process arguments into broadly accessible tickets.

Because the ephemeral-container record cannot be removed from the Pod spec, deleting the debugger process is not the same as restoring the original object. If the Pod is managed by a Deployment, StatefulSet, or another controller, prefer replacing the Pod through its normal controller after evidence is collected, following the service’s disruption policy. Do not edit the controller template to add an ephemeral container; it is a per-Pod troubleshooting action, not an application deployment change.

When a copied Pod is safer

If debugging requires changing the command, image, security context, or volume layout, kubectl debug --copy-to creates a separate Pod for investigation. That can be safer for repeatable reproduction, but the copy is not identical in every runtime or controller detail. Check which labels, service-account permissions, volumes, network policy, and secrets are inherited before starting it. A copy can accidentally remain routable or retain production credentials.

Ephemeral containers are also not a substitute for node debugging. If the failure is in the kernel, CNI, storage driver, or node runtime, a Pod-level debug container may not have the visibility or privilege required. Escalate to a dedicated node-debug workflow with explicit authorization rather than silently granting the in-Pod container host-level access.

Acceptance checks

Before closing the incident, confirm that the debugger used an approved immutable image, the expected target container and namespace were selected, the commands were read-only unless a change was approved, sensitive output was handled correctly, and the affected Pod was replaced or retained intentionally. Capture the cluster and kubectl versions because supported profiles and runtime behavior evolve.

Ephemeral containers make minimal production images diagnosable without rebuilding them during an incident. Their constraints are part of that safety model: they are observable, non-restartable, and not removable in place. Use them for bounded inspection, then return workload ownership to the controller and deployment process.

Related:

Sources:

Comments