Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

Kubernetes Pod Overhead and RuntimeClass: Account for the Runtime Around Containers

Model per-Pod sandbox overhead with RuntimeClass so Kubernetes scheduling, quota accounting, cgroups, and eviction decisions include real runtime costs.

A Pod consumes resources beyond the requests of its application containers. A sandboxed runtime may keep a virtual-machine monitor and guest kernel for each Pod; another runtime may have a smaller but still non-zero per-sandbox cost. If Kubernetes schedules only the container requests, a dense node can look safe on paper while the runtime infrastructure pushes the host toward pressure. Pod Overhead makes that per-Pod cost part of Kubernetes resource accounting, and RuntimeClass associates the cost with the runtime that incurs it.

This mechanism is deliberately narrower than capacity planning. It accounts for a configured, fixed amount per Pod; it does not measure the runtime process, reserve a separate host pool, or automatically discover a correct value. Operators must measure the runtime implementation, configure the appropriate class, ensure workloads select it, and verify that node capacity and quotas include the overhead.

RuntimeClass selects the execution contract

RuntimeClass is a cluster-scoped object that names a container runtime handler known to the node’s CRI implementation. A Pod opts into that class through spec.runtimeClassName. The class can also declare scheduling constraints so that a Pod is placed only on nodes that support the handler. Without a class selection, the Pod uses the cluster’s default runtime and does not receive overhead from an unrelated RuntimeClass.

The handler string is not an instruction to install a runtime. It must match configuration on the nodes’ container runtime, and every eligible node must support the handler. If runtime availability differs by node pool, configure a node selector or equivalent placement policy through the class and confirm that the selector is compatible with workload selectors. A class without a scheduling restriction is treated as available everywhere, so an inaccurate assumption can turn a valid API object into a Pod that fails after placement.

Pod Overhead has been stable since Kubernetes 1.24. The RuntimeClass admission controller copies the class’s overhead into the admitted Pod when the Pod references that class. Users should not set spec.overhead themselves: the controller manages that field, and a Pod that already supplies it is rejected. Verify the admitted object rather than assuming the manifest alone describes what the scheduler will evaluate.

Declare a measured fixed cost

The example below is schematic. kata-fc and its handler must correspond to a runtime installed and configured in the cluster; the resource quantities must come from measurements for that runtime and platform, not from this example.

apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
  name: sandboxed
handler: kata-fc
overhead:
  podFixed:
    cpu: 250m
    memory: 120Mi

A workload selects the class and continues to declare resources for its own containers:

apiVersion: v1
kind: Pod
metadata:
  name: isolated-worker
  namespace: batch
spec:
  runtimeClassName: sandboxed
  restartPolicy: Never
  containers:
    - name: worker
      image: registry.example.invalid/batch/worker:1.4.2
      resources:
        requests:
          cpu: "1"
          memory: 1Gi

The image address is intentionally a placeholder: replace it with an approved image in the deployment environment. The workload request describes the worker, while the RuntimeClass overhead describes the fixed per-Pod runtime cost. Neither value should be inflated to compensate for unrelated host services or uncertain application demand.

Understand what Kubernetes counts

For scheduling, the Pod’s overhead is added to the sum of its container requests. In this simple example, the scheduler accounts for 1.25 CPU and 1,144 MiB of memory (1 GiB plus 120 MiB), rather than only the worker’s 1 CPU and 1 GiB. That additional amount affects whether a node has enough allocatable resources for placement. With multiple containers, init containers, or restartable sidecars, use Kubernetes’ effective Pod request calculation for the target release rather than manually adding only the steady-state application containers.

ResourceQuota accounting also includes the overhead field along with container requests. That means a namespace can reject a Pod that would fit its quota based on application requests alone. This is expected: the sandbox consumes real namespace-attributed capacity. Quotas and node capacity are separate constraints, so an admitted Pod may still be unschedulable, and a free node does not override a namespace quota rejection.

After placement, the kubelet includes overhead when sizing the Pod cgroup. When every container defines a limit for a given resource, the Pod-level cgroup limit for that resource is based on the sum of container limits plus overhead. For Burstable and Guaranteed Pods, CPU shares are based on container requests plus overhead. Overhead also participates in Pod eviction ranking. It is therefore not merely a scheduler hint, nor is it a process-level limit that independently caps the runtime’s host-wide memory use. Node allocatable resources, kubelet reservations, cgroup configuration, and actual runtime behavior remain part of the operational model.

Keep the value honest and the class usable

Measure representative Pods on each runtime and node configuration. Include persistent sandbox processes and guest-kernel baseline memory, compare idle and loaded behavior, and understand whether the cost changes with architecture, runtime version, or enabled features. Because podFixed is a static value for the class, choose a conservative value that reflects the runtime contract without pretending that every Pod has identical instantaneous usage. If distinct configurations have materially different fixed costs, define distinct runtime classes rather than hiding that difference in one average.

The class is effective only when all of these conditions hold:

  • The RuntimeClass admission controller is enabled in the API server; otherwise the runtime overhead mutation is not performed.
  • The handler is configured in the CRI runtime on every node that can receive the Pod.
  • Scheduling constraints select only compatible nodes, including during node-pool upgrades and autoscaler replacement.
  • Workload templates explicitly set runtimeClassName; ordinary Pods do not inherit a non-default class automatically.
  • Namespace quota budgets include the overhead expected for the number and mix of selected Pods.
  • Monitoring and capacity reviews distinguish application requests from the added per-Pod cost.

Use the API to inspect the admitted value and node accounting:

kubectl get pod isolated-worker -n batch -o jsonpath='{.spec.runtimeClassName}{"\n"}{.spec.overhead}{"\n"}'
kubectl describe pod isolated-worker -n batch
kubectl describe node worker-7
kubectl get runtimeclass sandboxed -o yaml

The Pod’s .spec.overhead should match the intended class. In node descriptions, compare allocated requests with and without the class using a controlled test workload. For a full rollout, also inspect quota used values and scheduler events. If the admitted Pod has no overhead, first check runtimeClassName, the class object, and admission-controller configuration; if it has overhead but cannot be placed, compare the increased request with node allocatable capacity and placement constraints before lowering the value.

Treat RuntimeClass changes as scheduling changes

Changing overhead changes accounting for newly admitted Pods; it does not retroactively rebalance already scheduled workloads. Roll out an updated template through the owning controller, observe a representative replacement, and verify the resulting Pod specification, node allocations, and quota use. A higher overhead can make a rollout stall because replacement Pods no longer fit, especially when Deployment surge capacity is tight. Plan temporary quota and node headroom, while retaining realistic requests; reducing overhead merely to make a rollout pass defeats the accounting contract.

Also test upgrades of the runtime and node pool. The class name may remain constant while the handler implementation or its resource profile changes. Re-measure when the sandbox baseline changes materially, update the RuntimeClass deliberately, and canary the resulting Pod accounting before broad rollout. For incident response, preserve the admitted Pod spec and node allocation evidence so that later analysis can distinguish runtime overhead from application growth.

Pod Overhead is most useful when a runtime has a repeatable per-Pod infrastructure cost that would otherwise be invisible to the scheduler and quota controller. It makes that cost explicit across admission, placement, accounting, and eviction decisions. It does not replace correct container requests, node reservations, runtime configuration, or measured capacity planning.

Related:

Sources:

Comments