Kubernetes Local Ephemeral Storage: Accounting, Limits, and DiskPressure
Understand how Kubernetes accounts for emptyDir data, container logs, and writable layers, then size requests, limits, and node-level disk headroom.
Kubernetes local ephemeral storage is node-local capacity for scratch data, caches, container logs, and writable container layers. It is not a PersistentVolumeClaim: a Pod can be evicted, replaced elsewhere, or lost with its node, and local storage does not promise durability or a performance SLA. The operational challenge is accounting. A workload can write through several filesystem paths, while the kubelet’s ability to measure and enforce its declared budget depends on the node’s storage layout.
Know what the budget counts
With local ephemeral storage capacity isolation configured correctly, kubelet accounting includes disk-backed emptyDir volumes, container log directories, and writable container layers. Read-only image layers consume node disk too, but are shared runtime data rather than ordinary per-Pod writable-layer usage. Node filesystem pressure can therefore be caused by cached images or host data even when each Pod appears within its own budget.
A tmpfs volume, configured as emptyDir.medium: Memory, is charged as memory use, not local ephemeral storage. It belongs in the memory-risk review, not the disk budget. CSI ephemeral volumes are another distinct boundary: their storage is managed by a CSI driver and is not covered by the kubelet’s local ephemeral-storage Pod usage limits. Check the driver and cluster’s monitoring separately.
Requests reserve scheduling capacity; limits lead to eviction
Set resources.requests.ephemeral-storage from measured normal use plus expected growth. The scheduler uses requests against the node’s allocatable ephemeral storage; a request is not a reservation of physical blocks and does not stop a container from writing above it while the node has room. Limits are also per container. Kubelet can mark a Pod for eviction when an individual container’s writable layer and logs exceed its limit. For overall Pod accounting, it sums container limits and compares them with the combined usage of container files, logs, and disk-backed emptyDir volumes.
That response is eviction, not a synchronous per-container disk quota that makes a write fail exactly at the declared byte count. The kubelet measures usage on a schedule, so bursts can temporarily overshoot. A configured emptyDir.sizeLimit expresses a volume capacity bound; it does not reserve that capacity, and competing use of the node filesystem can make the volume run out of room sooner. Do not count the volume limit on top of the Pod limit: its usage contributes to the Pod’s total.
This example gives the application and a helper separate requests while keeping a shared scratch volume inside the Pod-wide total. The image names are placeholders; replace them with approved, digest-pinned images before deployment.
apiVersion: v1
kind: Pod
metadata:
name: api-example
namespace: production
spec:
containers:
- name: api
image: example.invalid/team/api:replace-me
resources:
requests:
ephemeral-storage: 1Gi
limits:
ephemeral-storage: 2Gi
volumeMounts:
- name: scratch
mountPath: /scratch
- name: log-helper
image: example.invalid/team/log-helper:replace-me
resources:
requests:
ephemeral-storage: 256Mi
limits:
ephemeral-storage: 512Mi
volumeMounts:
- name: scratch
mountPath: /scratch
volumes:
- name: scratch
emptyDir:
sizeLimit: 700Mi
Here the aggregate request is 1.25Gi and the aggregate limit is 2.5Gi. The 700Mi scratch ceiling is part of, not additional to, the aggregate usage. Container stdout and stderr are stored as node-level Pod logs and count too; prefer structured logs to stdout over writing a second unbounded log copy into emptyDir.
Verify metering before trusting a manifest
Kubelet measures storage only for supported filesystem layouts. In the common layout, /var/lib/kubelet and /var/log are on the node root filesystem. Moving kubelet or runtime data onto extra mounts can make usage reporting or limit enforcement incomplete unless the layout is one Kubernetes supports. Confirm the filesystem topology and kubelet’s reported capacity on every node pool; a valid resource field alone does not prove that the node measures it correctly.
The documented scanner periodically walks volume, log, and writable-layer directories. It can miss disk blocks held by an open file that has been unlinked. Project-quota monitoring can produce faster, more accurate measurements for supported filesystems, but it is an optional configuration, not a hard quota: current Kubernetes documentation marks LocalStorageCapacityIsolationFSQuotaMonitoring beta and disabled by default, and project quotas monitor usage rather than enforce it. Validate feature gates, user-namespace/runtime support, filesystem options, and version-specific documentation before enabling it.
To inspect the node’s advertised capacity and the kubelet Summary API for a specific running Pod:
NODE=$(kubectl get pod api-example -n production -o jsonpath='{.spec.nodeName}')
kubectl get node "$NODE" -o jsonpath='{.status.allocatable.ephemeral-storage}{"\n"}'
kubectl get --raw "/api/v1/nodes/$NODE/proxy/stats/summary" |
jq '.pods[] | select(.podRef.namespace == "production" and .podRef.name == "api-example") |
{pod: .podRef, totalBytes: ."ephemeral-storage".usedBytes,
containers: [.containers[] | {name, rootfsBytes: .rootfs.usedBytes, logBytes: .logs.usedBytes}]}'
Use this as a signal, not a complete node-disk inventory. Compare samples across startup, normal load, log bursts, cache warming, and rolling replacement. Also inspect node filesystem free bytes and inodes, kubelet events, container-runtime storage, and host-level consumers. Metrics Server’s familiar CPU and memory view is not a substitute for this storage evidence.
Treat DiskPressure as a separate failure path
A Pod that exceeds its measured local storage limit can be evicted even if the node has free disk. Separately, when a node filesystem crosses a kubelet eviction threshold, the node can report DiskPressure; kubelet first attempts node-level reclamation, such as garbage-collecting dead containers or deleting unused images, then may evict Pods. QoS class is not the eviction-order key for ephemeral-storage pressure. Selection considers whether a Pod’s usage exceeds its requests, its Priority, and its usage relative to requests. These are node-level recovery decisions, not PVC resize events.
For namespace quota, an administrator must configure an ephemeral-storage ResourceQuota and workloads need declared ephemeral-storage limits for that quota to be enforced. In production, test both paths on a disposable node: exceed one Pod’s budget with an isolated scratch writer, then separately trigger a controlled node-pressure threshold. Confirm the expected eviction event, replacement behavior, data loss from emptyDir, alerting, and recovery. Never use real node exhaustion as a first test, and keep durable state on storage designed to survive Pod replacement.
Related:
- Kubernetes QoS Classes: How Requests and Limits Shape Eviction and Scheduling
- Kubernetes PVC Expansion: Growing Persistent Volumes Without Data Loss
Sources: