Kubernetes CPU Throttling: Diagnose Quota Pressure and Latency
Diagnose Kubernetes CPU throttling with cgroup and Prometheus evidence, separate requests from limits, and tune quotas without masking capacity issues.
A Pod can be healthy, use less CPU than its node appears to have available, and still show high request latency because one of its containers is being throttled by a CPU limit. Kubernetes CPU requests and limits answer different questions: requests influence scheduling and relative CPU allocation under contention; limits impose a hard ceiling on CPU time. The distinction is easy to miss when dashboards display only current CPU usage or when teams raise requests without changing the restrictive limit.
CPU throttling is not automatically a fault. It is the kernel enforcing configured bandwidth. The operational question is whether the enforcement is coincident with an application SLO regression, queue growth, timeouts, or long tail latency. This guide connects Kubernetes configuration to Linux cgroup evidence and cAdvisor metrics so a team can establish that causal link before tuning a workload.
Separate requests, limits, and observed usage
Kubernetes schedules based on resource requests. A container can use more CPU than its request when its node has available capacity, and requests also influence its relative share when CPU is contended. A CPU limit is different: on Linux, the kubelet and container runtime configure kernel control groups so the workload cannot continuously consume more than its limit. When the workload uses its allowed CPU quota for an enforcement interval, runnable tasks may wait for a later interval even if another CPU on the node would otherwise be idle.
This can produce a counterintuitive symptom: the node is not saturated, the container’s measured CPU usage is near its configured ceiling, but latency spikes during bursts. The kernel is enforcing the container’s quota, not allocating a guaranteed amount of CPU independently of that quota. A higher request alone does not remove the ceiling. Conversely, removing a limit does not give the Pod a reservation or guarantee latency when the node is contended.
Before diagnosing throttling, inspect the effective resources after admission. A namespace LimitRange may default requests or limits, and Kubernetes can use a specified limit as the request when no request was supplied and no admission mechanism added one. Check the live Pod rather than only the Deployment source:
kubectl get pod api-7d86b6c9d8-abcde -n production -o yaml
kubectl describe pod api-7d86b6c9d8-abcde -n production
kubectl get limitrange,resourcequota -n production
Review every container in the Pod, including proxies, telemetry agents, and sidecars. A small per-container limit can be reached by a multi-threaded process during a short burst, and the main application may not be the only container affected when a request path depends on a throttled sidecar.
How cgroup CPU bandwidth limits behave
On Linux, container runtimes normally express CPU resource controls through cgroups. A quota is an aggregate time budget over a period, not a promise that a fixed CPU core is reserved for the container. The cgroup v2 cpu.max file exposes the maximum bandwidth in the form MAX PERIOD; max means no bandwidth cap. The actual path and cgroup layout depend on the node’s runtime and system manager, so do not hard-code a host path copied from a different distribution.
The cgroup v2 cpu.stat file includes nr_periods, nr_throttled, and throttled_usec when the CPU controller is enabled. These cumulative counters make it possible to compare throttling over the same interval as application latency and traffic. They account for throttling caused by that cgroup’s own bandwidth limit; parent-cgroup constraints can require examining ancestor limits and the kernel’s local statistics too. A container-level graph can therefore be incomplete if a Pod or higher-level cgroup also has a restrictive quota.
Kubernetes cgroup behavior also depends on the node release, kernel, runtime, and cgroup version. Current Kubernetes documentation marks cgroup v2 stable since Kubernetes v1.25 and cgroup v1 deprecated since v1.35; newer kubelets do not start on cgroup v1 by default unless an administrator changes the documented policy. For current clusters, verify the cgroup version and runtime rather than assuming a cgroup v1 file layout.
On a node where cgroup v2 is in use, an authorized operator can inspect the relevant control group from the host:
cat /proc/$(pgrep -f 'your-process-pattern' | head -n 1)/cgroup
CGROUP_DIR='/sys/fs/cgroup/REPLACE_WITH_RESOLVED_CGROUP_PATH' # Set this to the path resolved from the process cgroup entry.
cat "$CGROUP_DIR/cpu.max"
cat "$CGROUP_DIR/cpu.stat"
These commands are templates, not a universal container lookup procedure. A process match may be ambiguous, the cgroup path is runtime-specific, and a process can exit between commands. Resolve the container and its cgroup using the node’s supported runtime tooling before reading files. Do not grant an application container broad host access just to scrape host cgroups; use the cluster’s approved node-level metrics pipeline.
Measure throttling with correctly interpreted counters
cAdvisor exposes cumulative counters including container_cpu_cfs_periods_total, container_cpu_cfs_throttled_periods_total, and container_cpu_cfs_throttled_seconds_total. The first counts enforcement intervals; the second counts intervals in which tasks were throttled; the third reports cumulative throttled time. These metrics are emitted only where the applicable CPU quota statistics are available, and metric names or label sets can differ across collection pipelines. Confirm that the series exists in your own Prometheus before using it in an alert.
A useful starting ratio is the fraction of observed enforcement intervals that included throttling:
sum by (namespace, pod, container) (
rate(container_cpu_cfs_throttled_periods_total[5m])
)
/
clamp_min(
sum by (namespace, pod, container) (
rate(container_cpu_cfs_periods_total[5m])
),
1e-9
)
This query assumes the scrape pipeline attaches namespace, pod, and container labels consistently. Some cAdvisor deployments expose only an id or runtime-specific labels, so inspect /metrics and configure relabeling before copying the grouping verbatim. rate() is appropriate for these counters; do not graph raw lifetime totals as if they were current pressure.
The ratio is not the percentage of wall-clock time lost, nor does it prove that a user request waited. It records periods with throttling, not how important those periods were to the application’s critical path. Pair it with request duration percentiles, queue depth, concurrency, CPU usage relative to limit, restart/error rates, and node-level saturation. A steadily increasing throttled-time counter with flat latency may be acceptable; a modest ratio aligned with a sharp p99 regression during bursts may warrant investigation.
Avoid alerting on a single brief spike. Use a sustained window and a workload-level SLO condition, then drill down to the affected container. A shared query aggregated across all namespaces can hide the one sidecar or replica that is throttled, while a high-cardinality alert for every short-lived Pod can create its own operational load.
Prove or reject CPU throttling as the cause
Use evidence from the same time interval rather than treating a metric as a root cause by itself:
- Correlate with the service symptom. Compare throttling rates with request latency, timeouts, queue length, and traffic. If the service is slow while throttling remains zero, investigate other bottlenecks.
- Check the effective limit. Inspect the live Pod, admission defaults, and all containers. Confirm the measurement unit: Kubernetes CPU quantities such as
500mrepresent half a CPU core. - Compare usage and quota. A container continuously near its limit and showing throttled periods is stronger evidence than a single counter increase. Usage averaged over five minutes can hide short bursts, so compare intervals suitable for the workload’s latency target.
- Separate local quota from host contention. Inspect node CPU pressure, runnable queue, steal time in virtualized environments, and ancestor cgroup limits. A throttled child may not tell the whole hierarchy story.
- Test a controlled change. In a representative environment or canary, adjust only the limit or burst policy under review. Preserve the request, replicas, traffic, and application version where possible, then compare p95/p99, CPU consumption, throttling, and node pressure.
Do not assume that raising a CPU request will help when the workload is already bound by its own limit. Do not remove a limit across a fleet as the first experiment: an unbounded CPU consumer can starve neighbors under host contention. Decide whether the limit is a tenant safety boundary, a cost control, or a latency policy, and define what tradeoff is acceptable before changing it.
Tune the workload and the policy together
For a latency-sensitive service, set a request that reflects realistic sustained demand and the capacity the scheduler must reserve. Then choose a limit based on measured burst behavior, node density, and the cluster’s isolation model. In some clusters, platform policy intentionally omits CPU limits for selected workloads while retaining requests; in others, limits are mandatory guardrails. Neither approach is universally correct. The workload’s measured CPU profile and the consequences of noisy neighbors determine the tradeoff.
When a limit is too low, increasing it can reduce throttling but may let the workload consume more node capacity during simultaneous bursts. Scaling out can lower per-replica CPU demand, but it increases scheduler, connection, and downstream load. Optimizing hot paths, reducing lock contention, batching work, and bounding application concurrency may improve latency without increasing the cluster’s CPU budget. Recheck capacity and autoscaler behavior after any change.
If container startup configuration changes CPU resources, use a rollout or a supported in-place resize path only after confirming the cluster version, feature behavior, and workload policy. Do not edit cgroup files manually as a persistent fix: the kubelet and runtime reconcile settings from the Pod specification and can overwrite ad-hoc changes.
Common misdiagnoses
| Observation | Why it is not enough | What to correlate |
|---|---|---|
| Node CPU is below 100% | A per-container quota can throttle while capacity exists elsewhere on the node | Container quota, cgroup counters, and parent limits |
| Container CPU usage is below its limit | A long-window average can hide brief bursts that exhaust a quota | Shorter rates, application latency, and concurrency |
| Throttled time increased | Counters are cumulative; an increase alone does not show current impact | rate() over the same period as SLO metrics |
| Pod restarted | CPU throttling does not itself normally terminate a container | Exit reason, events, OOM evidence, and liveness failures |
| Request was raised | The container may still have the same hard CPU limit | Effective live limit and resulting cgroup quota |
| Metric is absent | Scrape configuration, cAdvisor version, label relabeling, or no CPU quota may explain it | Raw exporter endpoint and cgroup statistics on the node |
Production review checklist
- Live requests and limits are understood after
LimitRange, policy, and admission defaults. - Every container on the request path is included, not just the main application container.
- CPU throttling metrics are present, correctly labeled, and queried as counter rates.
- Throttling is correlated with workload SLOs and concurrency before a resource change is proposed.
- Node and ancestor-cgroup pressure are checked alongside the container’s own quota.
- A limit increase, omission, or scaling change is canaried with capacity and tenant-isolation effects measured.
- The cgroup version, kernel, runtime, and kubelet behavior match the node’s current supported configuration.
- The final request/limit policy is documented as a deliberate latency, reservation, and isolation tradeoff.
CPU throttling is a mechanism, not a diagnosis. Kubernetes configures a resource policy, the runtime translates it into cgroup controls, and Linux enforces the CPU bandwidth. Use workload symptoms plus cgroup and exporter evidence to identify which boundary is binding; then change one part of the policy at a time and verify both latency and cluster capacity.
Related:
- Kubernetes Vertical Pod Autoscaler: Recommendation, Update Modes, and Safe Rollouts
- Kubernetes Node Allocatable: Reservations, Eviction Headroom, and Real Capacity
Sources: