Kubernetes CPU and NUMA Management: Exclusive Cores Without Surprises
Configure Kubernetes static CPU management and NUMA alignment, understand Guaranteed Pod eligibility, and prevent topology admission failures in production.
Kubernetes’ default CPU allocation is appropriate for most services: the scheduler accounts for CPU requests, and the runtime enforces CPU limits, but containers may share CPUs on a node. Latency-sensitive workloads sometimes need stronger placement properties, such as exclusive CPUs or CPU, memory, and device allocations that share a NUMA locality. The kubelet’s CPU Manager and Topology Manager provide node-local controls for those cases.
These controls are not a cluster-wide reservation service and do not make a Pod topology-aware to the scheduler. They change what can happen after a Pod lands on a node: its containers may receive an exclusive CPU set, a topology hint, or a node-local admission rejection. Use them only after measuring a workload that benefits from CPU cache affinity, reduced contention, or device locality, and plan for node-specific fragmentation and recovery.
Separate CPU shares, exclusive CPUs, and NUMA alignment
Three behaviors are commonly confused:
- CPU requests and limits describe scheduler accounting and runtime enforcement. They do not, by themselves, pin a container to exclusive logical CPUs.
- CPU Manager static policy can assign exclusive CPUs to eligible containers on one node. It does not tell the cluster scheduler which node has a useful CPU set.
- Topology Manager coordinates locality hints from kubelet resource managers and device plugins. It can prefer or require a NUMA-compatible allocation, but it does not reserve all resources in the cluster or make
kube-schedulertopology-aware.
The word “exclusive” is relative to the kubelet-managed CPU sets. Correct reservation of CPUs for the operating system and Kubernetes daemons still matters, and the physical CPU topology and hardware threading configuration affect how logical CPU IDs map to cores. Validate the actual node’s topology and container runtime behavior instead of assuming that CPU IDs 0-3 represent four isolated physical cores.
CPU Manager’s static policy
With the default none policy, containers use the node’s shared CPU pool, subject to their cgroup quota and weight. The static policy keeps a shared pool while allowing eligible containers to receive exclusive CPU sets. Under the documented rule, only a container that belongs to a Guaranteed Pod and has an integer CPU request can receive exclusive CPUs. BestEffort and Burstable containers, and Guaranteed containers with fractional CPU requests, remain in the shared pool.
Guaranteed QoS is stricter than simply setting one CPU limit. For the Pod to be Guaranteed, all relevant containers must have CPU and memory requests and limits, and each request must equal its limit. CPU Manager’s eligibility is then evaluated per container: a 2 CPU request can qualify, whereas 1500m does not qualify for exclusive assignment. Check the resulting .status.qosClass and each container’s effective resource values; do not infer QoS from a Deployment-level template fragment alone.
Static policy needs a nonzero CPU reservation for kubelet and system use. Reserved CPUs are not part of the exclusive workload pool. If the reservation is undersized, system processes may compete with workloads; if it is too large, fewer CPUs remain for allocatable Pods. Align kube-reserved, system-reserved, explicit reserved CPU lists, scheduler allocatable capacity, and workload requests. CPU reservations and CPU Manager policy are per-node kubelet configuration, so node pools with different hardware should be configured and tested independently.
The policy maintains a checkpoint, normally at /var/lib/kubelet/cpu_manager_state. Changing from none to static is not just a live toggle: the checkpoint records the old policy and assignments. The upstream procedure drains the node, stops kubelet, removes the stale CPU Manager checkpoint, changes the policy, and restarts kubelet. Apply the distribution-supported procedure, verify the kubelet root directory, and never delete that state file on a live node as an exploratory fix. A changed set of online CPUs can also require a drain and state reset.
Topology Manager coordinates hints, not arbitrary placement
CPU Manager, Memory Manager, and device plugins can provide hints describing NUMA nodes on which requested resources are available. Topology Manager combines those hints under a kubelet policy. Its scope controls whether it considers each container separately (container, the default) or groups a Pod’s containers (pod). Its policy controls whether it merely records a preference or requires a preferred locality.
| Policy | Admission behavior | Operational trade-off |
|---|---|---|
none |
No topology alignment | Broadest admission; locality is left to individual resource managers |
best-effort |
Records the preferred hint and admits even if the hint is not preferred | Preserves availability but may not deliver the requested locality |
restricted |
Rejects admission when the best hint is not preferred | Protects the locality objective but can leave work requiring replacement |
single-numa-node |
Requires a suitable allocation on one NUMA node | Strongest locality constraint; can sharply reduce the nodes that can run the Pod |
With container scope, different containers in the same Pod can be aligned independently. Pod scope asks the kubelet to place all containers within a common NUMA-node set; combining Pod scope with single-numa-node can keep a latency-sensitive, tightly coupled workload on one NUMA node when all required resources fit there. That constraint is meaningful only when the relevant resource managers and device plugins provide topology hints.
The scheduler does not account for the final Topology Manager admission decision. A Pod can bind to a node and then be rejected locally because the CPU, memory, and device hints cannot satisfy the selected policy. The scheduler will not automatically reschedule that already-terminated Pod; a Deployment or another controller can create a replacement, but repeated replacements may land on equally unsuitable nodes. Watch for topology admission errors, select an appropriate node pool, and test that your owning controller actually recovers.
Design requests for whole-core and local allocations
Start with an explicit workload requirement, not a policy name. If the workload needs protection from CPU sharing, verify it is Guaranteed with integer CPU requests. If it depends on a hardware accelerator, inspect that device plugin’s topology hints and confirm the desired CPU, device, and memory resources can coexist on the same NUMA node. If it requires low tail latency but can tolerate cross-node alignment, a preference-oriented policy may be more suitable than rejecting Pods.
For workloads that need full physical cores rather than individual SMT threads, current CPU Manager policy options include full-pcpus-only; the exact availability and combination rules depend on Kubernetes version and feature-gate configuration. Other options influence distribution across NUMA nodes or CPU cores. Do not copy a policy option from current upstream documentation into an older cluster without checking its version and gate requirements. Incompatible options can prevent kubelet startup, and changing some options requires resetting CPU Manager state through a controlled node procedure.
Topology Manager also has node-count and performance boundaries. The current upstream documentation describes a default limit of eight NUMA nodes for Topology Manager, reflecting the cost of evaluating many possible affinities. A max-allowable-numa-nodes option exists in newer Kubernetes versions, but upstream cautions that it is not recommended for nodes above eight NUMA nodes due to limited impact data. High-socket-count systems need release-specific validation; do not assume the highest configurable value is a production recommendation.
The scheduler’s resource fit remains necessary but insufficient. Reserve enough CPUs for daemons, keep allocatable CPU consistent with those reservations, and ensure the node pool contains enough whole-core and NUMA-local capacity for the intended Pod sizes. A cluster may have abundant aggregate free CPU while no single NUMA node has the locality required by single-numa-node. Conversely, a relaxed policy may admit the Pod but provide no performance benefit if the resource providers’ hints are absent or poor.
Roll out as a node-pool change
Before changing kubelet policy, capture the node CPU topology, NUMA layout, CPU reservations, kubelet version and configuration, container runtime, active CPU Manager checkpoint, allocatable resources, and representative workload performance. Establish a baseline for throughput, tail latency, CPU throttling, and accelerator locality. A single successful Pod placement is not proof that the target node pool has enough usable capacity under rollout or failure conditions.
For a static-policy transition, cordon and drain one canary node using the distribution-supported procedure, stop or reconfigure kubelet only as directed, reset the checkpoint if required, and return the node to service. Confirm that existing and newly created Pods behave as expected, that system daemons retain CPU capacity, and that the node’s allocatable values match the intended reservation. Continue pool by pool, leaving enough healthy capacity to handle workload replacement while each node is unavailable.
Test an eligible Guaranteed Pod with integer CPU requests, a fractional-CPU Guaranteed Pod, a Burstable Pod, and a topology-constrained Pod that deliberately cannot fit. For the last case, validate both the kubelet’s rejection reason and the owning controller’s replacement behavior. If accelerators are involved, test allocation under resource contention and verify that the CPU and device assignments match the locality goal; kubectl describe alone may not expose all host-level cpuset details.
Observe the node-local outcome
Inspect the actual kubelet configuration and node allocatable capacity, then compare it with Pod admission events and QoS status:
kubectl describe node worker-accelerator-03
kubectl get pod -n inference model-worker-0 -o jsonpath='{.status.qosClass}{"\n"}'
kubectl describe pod -n inference model-worker-0
kubectl get events -n inference --sort-by=.lastTimestamp
On a self-managed node, inspect the CPU Manager checkpoint and the runtime’s effective CPU set using the distribution’s supported host-debugging procedure. Cgroup paths and interfaces vary by runtime and cgroup version, so avoid hard-coding a host path into an application diagnostic. Correlate the observed CPU set with lscpu or the node’s NUMA inventory rather than assuming a list of logical CPU IDs means one full physical core each.
Monitor Pending and terminated Pods, Topology Affinity admission errors, CPU throttling, shared-pool saturation, kubelet restarts, and checkpoint restore errors. If the node’s online CPU topology changes after firmware updates or hardware replacement, repeat the inventory and follow the documented drain and checkpoint reconciliation path before scheduling production workloads there.
Production acceptance checklist
- The workload has a measured reason to require exclusive CPUs or NUMA locality.
- Every eligible container meets Guaranteed QoS and integer CPU request requirements.
- CPU reservations leave a functioning shared pool and sufficient capacity for system services.
- Kubelet and resource-manager configuration is consistent across like-for-like nodes and version-gated options are verified.
- Topology scope and policy match whether the workload needs a preference or a hard admission constraint.
- The scheduler’s lack of Topology Manager awareness is accounted for in node selection and replacement behavior.
- Canary rollouts, node drains, CPU hotplug or replacement, and checkpoint reset procedures are tested.
- Performance is measured against a baseline under representative contention, not inferred from a successful Pod start.
CPU Manager and Topology Manager are specialized tools for the node-local part of performance engineering. They can reduce noisy-neighbor interference and coordinate CPU, memory, and device locality, but they also create stricter admission constraints and operational state on each kubelet. Production readiness means proving the resource fit and recovery path on the actual hardware pool, not merely enabling static and seeing a Pod run.
Related:
- Kubernetes QoS Classes: How Requests and Limits Shape Eviction and Scheduling
- Kubernetes Topology Spread Constraints: Design for Failure Domains
Sources: