Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

Kubernetes ResourceQuotas in Production: Admission Budgets and Failure Modes

How namespace quotas govern Kubernetes admission, compute and storage budgets, object counts, LimitRange defaults, and failure diagnosis.

Kubernetes ResourceQuota is a namespace-scoped admission budget. It can cap aggregate requests, limits, storage claims, and selected API object counts. It does not reserve node capacity, meter application CPU or memory at runtime, or evict workloads when a limit is reduced. Those boundaries determine whether a quota is a useful guardrail or a source of confusing rollout failures.

Quota is checked at the API boundary

The ResourceQuota admission controller evaluates creates and updates against the quota objects that apply to the namespace. If accepting a request would exceed a hard limit, the API server rejects that request with 403 Forbidden and identifies the quota and resource. The quota admission plugin must be enabled for these policies to be enforced; Kubernetes includes it among the recommended admission controllers, but managed-cluster operators should verify the control-plane configuration rather than infer it from the manifest.

A Deployment can be stored successfully even when its desired Pods cannot be admitted. The Deployment controller creates or updates a ReplicaSet; that ReplicaSet then attempts Pod creation, where quota admission can reject each excess Pod. The workload object existing in the API is therefore not proof that the requested capacity is available. Inspect controller events and admitted child objects.

Define what the budget measures

Resource quotas measure declared requests and limits as accounted by Kubernetes, not instantaneous process consumption. For example, requests.memory is a namespace sum of memory requests for Pods in non-terminal states; it is not a live memory meter or a memory reservation on a particular node. limits.memory sums declared limits; it does not make aggregate actual memory usage stop precisely at that number. The scheduler, kubelet, cgroups, and node-pressure eviction have separate responsibilities.

apiVersion: v1
kind: ResourceQuota
metadata:
  name: payments-budget
  namespace: payments
spec:
  hard:
    requests.cpu: "16"
    requests.memory: 32Gi
    limits.cpu: "32"
    limits.memory: 64Gi
    requests.storage: 2Ti
    persistentvolumeclaims: "40"
    pods: "250"
    count/jobs.batch: "500"
    services.loadbalancers: "2"

The compute keys aggregate resource declarations across admitted Pods. requests.storage caps the total storage requested by claims, while persistentvolumeclaims caps their count. A pods quota counts Pods in non-terminal phases. count/jobs.batch is an API-object count quota for Jobs; it is useful when a faulty or overly prolific CronJob could leave a large number of Job objects behind. Object-count limits protect API and controller capacity as well as application-facing resources.

Each hard key is an independent ceiling. Quotas are expressed in absolute quantities, not as a percentage of the cluster. Adding nodes does not automatically enlarge them, and setting namespace ceilings whose combined demand exceeds real cluster capacity does not create that capacity. Quotas also do not constrain placement to a team’s nodes: Pods from different namespaces may share a node. Capacity policy must account separately for node allocatable resources, system workloads, disruption headroom, and scheduling constraints.

Storage budgets can be partitioned by StorageClass when teams use distinct service tiers. For example, <class>.storageclass.storage.k8s.io/requests.storage limits the sum of PVC requests for that class, while the corresponding persistentvolumeclaims key caps its claim count. This constrains requested capacity in the API; it is not a measurement of backend bytes consumed after filesystem overhead, snapshots, or provider-side copies.

Pair aggregate quotas with per-object defaults

When a namespace quota covers CPU or memory, new Pods need the corresponding request or limit values required by that quota. Omitting them can cause Pod admission to fail, including when a controller creates the Pod on behalf of a Deployment. A LimitRange can supply defaults and bound each individual container or Pod, while ResourceQuota constrains the namespace total. They solve different problems.

apiVersion: v1
kind: LimitRange
metadata:
  name: container-sizing
  namespace: payments
spec:
  limits:
    - type: Container
      defaultRequest:
        cpu: 250m
        memory: 256Mi
      default:
        cpu: "1"
        memory: 1Gi
      min:
        cpu: 100m
        memory: 128Mi
      max:
        cpu: "4"
        memory: 8Gi

Defaults are injected during Pod admission; they do not rewrite existing Pods. Check the final admitted Pod specification, because sidecars and other admission mutations also consume quota. Keep a single, intentional source of defaults in a namespace: if multiple LimitRange objects provide defaults, Kubernetes does not guarantee which default is applied. Ensure every default limit is at least as large as its request. Kubernetes does not validate LimitRange defaults against values supplied by a client; for example, a default limit lower than a submitted request can make Pod admission fail validation. Test the resulting API request and handle an admission rejection separately from an admitted-but-unschedulable Pod.

Separate quota exhaustion from scheduling failure

These two failures occur at different points. A quota rejection means the API server refused to persist the requested child object. A scheduling failure means the Pod exists but the scheduler cannot place it on a node. kubectl get pods alone is not enough to distinguish them: a rejected Pod may never appear in the list, while an admitted but unschedulable Pod remains visible with scheduling events.

Use the namespace and controller evidence together:

kubectl get resourcequota -n payments
kubectl describe resourcequota payments-budget -n payments
kubectl describe deployment api -n payments
kubectl describe replicaset -n payments
kubectl get events -n payments --sort-by=.metadata.creationTimestamp

For a quota rejection, compare the used and hard entries and read the event or API error for the specific key, such as requests.memory. For a Pod in Pending, inspect its scheduling events and node allocatable resources instead of raising quota blindly. kubectl describe resourcequota shows the observed namespace usage alongside each hard value; retain this output with rollout diagnostics so an incident review can establish what the admission controller was enforcing at the time.

Budget for controller-driven peaks

Steady-state replica count is not necessarily the quota needed to release a workload. A Deployment rolling update can create surge Pods while old replicas still exist; maxSurge controls how many it may create above the desired count, and terminating Pods can keep consuming resources while replacements start. If quota has no room for the new Pod’s requests or count, the Deployment object can remain accepted while its rollout stalls at Pod admission. Reducing maxSurge trades some rollout speed for lower temporary demand; setting maxUnavailable changes the availability side of that tradeoff, so do not tune it as a quota workaround without validating service capacity.

An HPA’s maxReplicas is an autoscaling ceiling, not a quota exemption or capacity reservation. Ensure the namespace budget can admit the intended replica range and the rollout headroom for those replicas. Otherwise a scale decision can be written to the workload while one or more desired Pods are rejected by quota. Confirm the HPA’s conditions and desired/current replicas together with Deployment, ReplicaSet, and quota events before diagnosing a scaling incident.

Treat changes as forward-looking policy

Reducing a quota does not terminate, resize, or otherwise alter already-created objects. It changes which future creates and updates can be admitted. If the new hard value is below current usage, existing workloads can continue, but new Pods, claim growth, or other quota-consuming changes may be rejected until usage falls or the budget is revised. A rollout can therefore fail to replace old replicas even though the existing application remains healthy.

Quota is also not an automatic cluster reservation system. If several namespaces have ceilings that exceed aggregate capacity, scheduling contention is still resolved by the scheduler and the resources available on individual nodes. Conversely, a namespace may have unused quota while a Pod remains unschedulable because no eligible node has the required resources or topology. Keep per-namespace budgets, cluster capacity planning, and per-workload requests as separate inputs to the same release review.

Make quota policy testable

For each namespace, document the hard keys, the workload classes they cover, the intended capacity headroom, and the owner of budget changes. Validate the policy using a temporary namespace or a controlled test workload before relying on it in production. Include at least one request that should be admitted, one that should exceed each critical quota, and a controller-managed rollout that proves the resulting event is visible to operators.

An operational acceptance check should confirm that:

  • ResourceQuota objects are present in the intended namespaces and report the expected hard values.
  • New Pods receive deliberate requests and limits after all admission mutations, and stay within both per-object LimitRange constraints and namespace totals.
  • A deliberately oversized test request is rejected with an actionable quota error, while an in-budget request is admitted.
  • A Deployment whose child Pod is rejected is diagnosed from ReplicaSet or namespace events rather than mistaken for a scheduler failure.
  • Lowering a quota leaves existing objects untouched but blocks a measured over-budget create or update as expected.
  • Adding cluster capacity, changing node allocatable values, or changing a workload’s placement constraints is not incorrectly assumed to change the namespace quota.

The objective is predictable admission and fair, reviewable budgets - not a claim that a namespace is isolated from every other tenant or guaranteed its full quota under contention.

Related:

Sources:

Comments