Kubernetes Cluster Autoscaler: Pending Pods, Node Groups, and Safe Scale-Down
Diagnose Cluster Autoscaler decisions through scheduling fit, node-group templates, disruption budgets, scale-down blockers, cloud limits, and status events.
The Kubernetes Cluster Autoscaler (CA) changes the size of configured node groups when it decides that additional nodes could make unschedulable Pods schedulable, or that existing nodes are no longer needed and their movable Pods can run elsewhere. It is not a CPU-utilization controller and it does not create a workload replica. The scheduler, workload autoscaler, Cluster Autoscaler, cloud provider, and node bootstrap system each own a different part of the path.
This distinction is central to troubleshooting. A Pending Pod does not guarantee scale-up: Cluster Autoscaler must determine that a configured node group can produce a node satisfying that Pod’s requests, affinity, topology, taints, volume constraints, and other scheduling rules. A low-CPU node does not guarantee scale-down: the Pods on it must be safely movable, disruption budgets must permit eviction, and the node group must be above its minimum size. Successful cloud API requests also do not prove that a new machine registered as a healthy Kubernetes Node.
Understand the responsibility boundaries
The scheduler assigns Pods to nodes that currently exist. An HPA or another workload controller may increase replica count in response to metrics. If resulting Pods cannot schedule, Cluster Autoscaler evaluates whether adding a node from a configured node group would help. The cloud provider then provisions or resizes infrastructure, and the bootstrap process must install and register that node. After registration, the scheduler still has to bind the Pod and the kubelet must start it.
The scale-up chain therefore has multiple observable gates:
- A controller creates a Pod or increases replicas.
- The scheduler reports the Pod as unschedulable and records the reasons.
- Cluster Autoscaler recognizes the Pod and simulates it against configured node-group templates.
- If a group can help and constraints permit, Cluster Autoscaler requests a group resize.
- The provider creates a machine within quota and account limits.
- Node bootstrap, networking, kubelet, and node agents bring the machine into the cluster.
- The scheduler retries the Pod against the new Node.
If a Pod is Pending, collect scheduler events before changing autoscaler flags. If a resize request was issued but no Node appears, investigate provider quota, VM provisioning, image/bootstrap, network, credentials, node registration, and node health. Cluster Autoscaler does not own every later step in that chain.
Scale-up is about schedulability, not observed CPU
Cluster Autoscaler reacts to Pods the scheduler marks unschedulable, not to a CPU graph crossing a percentage. The Pod’s resource requests affect the scheduler’s fit decision. Actual low utilization does not mean a large-request Pod can fit, and actual high utilization alone does not create a new node if no unschedulable workload requires one.
Inspect the Pod’s events, requests, affinity, selectors, topology spread, taints/tolerations, PVC topology, and priority. Then determine which node groups the autoscaler is configured to manage and whether a hypothetical node from one of them would satisfy all relevant constraints. A node type with enough CPU but the wrong zone, architecture, GPU resource, label, or taint is not a useful scale-up option.
kubectl describe pod api-7b8c9f6c6d-abcde -n production
kubectl get pod api-7b8c9f6c6d-abcde -n production -o yaml
kubectl get nodes -L node-pool,topology.kubernetes.io/zone,kubernetes.io/arch
kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml
kubectl -n kube-system get events --sort-by=.lastTimestamp
Cluster Autoscaler exports decisions through logs, Pod and Node events, and a status ConfigMap in deployments that publish it. Look for scale-up or “did not trigger scale-up” evidence and correlate timestamps with scheduler events. A status ConfigMap that is not being updated may indicate that the autoscaler is unhealthy or unable to communicate with the API; check the actual deployment name and status resource configured for the installation.
Make node-group templates match real nodes
The autoscaler simulates a future node using node-group information. Groups should have consistent capacity, labels, and system workloads so a simulated node is representative. For scale-from-zero, provider-specific node-template metadata must describe the labels, taints, resources, and topology a future Node will receive. If the metadata is missing or inaccurate, the autoscaler may reject a Pod that would fit after boot, or choose a group that creates a Node on which the Pod cannot run.
Node-group auto-discovery and permissions are provider-specific. Follow the official integration guide for the installed cloud provider, scope the autoscaler’s cloud permissions to the intended groups, and explicitly manage minimum, maximum, and cluster-wide limits. A group at its maximum cannot satisfy additional demand; a group at its minimum may not scale down even when a particular node looks idle. Check provider account quotas before raising maxima, because a configuration limit is not a capacity reservation.
Do not manually resize or replace nodes inside a group managed by Cluster Autoscaler unless the provider’s documented workflow requires it. The autoscaler assumes the group state and nodes reflect the same source of truth. Running a second node-group autoscaler against the same groups can produce conflicting resize decisions; make ownership exclusive and documented.
Scale-down requires safe Pod movement
Cluster Autoscaler considers a node for removal only when it is under the configured utilization threshold for the required period and its Pods can be relocated. Its utilization calculation is based on requested CPU and memory relative to node allocatable, not live CPU usage alone. A node can show low actual CPU and still be ineligible because requests are large, its group is at minimum size, it was recently added, or some Pod cannot move.
Common blockers include restrictive PodDisruptionBudgets, standalone Pods without a controller to recreate them, local storage or topology constraints, affinity that leaves no destination, kube-system workloads without an acceptable disruption path, and Pods protected from eviction by autoscaler-specific settings. DaemonSet and mirror Pods receive special handling, with behavior controlled by the autoscaler version and configuration. Do not make a PDB permissive merely to remove one node until the application’s disruption objective and replica health are understood.
Scale-down can proceed through Kubernetes eviction and drain behavior. The autoscaler must respect applicable disruption constraints and wait for Pod termination within configured limits. A drain blocked by a PDB, a stuck terminating Pod, or a workload that cannot fit elsewhere can leave a node active despite low utilization. Check the events and logs for the specific Pod and node rather than repeatedly lowering the utilization threshold.
For a temporary maintenance exception, the project documents the node annotation cluster-autoscaler.kubernetes.io/scale-down-disabled=true to exclude a node from scale-down. Use it with an owner, reason, expiry or review date, and automation that removes it when the exception ends. An unowned permanent annotation silently consumes capacity and can mask a real autoscaling configuration issue.
Coordinate with workload scaling and topology
When HPA increases replicas, Cluster Autoscaler may add nodes only after Pods become unschedulable and only if a node group can satisfy them. End-to-end scale-out time includes metric collection, HPA reaction, scheduler decisions, autoscaler evaluation, provider provisioning, node bootstrap, image pulls, and application readiness. Increasing CA scan frequency cannot eliminate a slow VM launch, an unavailable image registry, or an application that takes several minutes to become Ready.
Topology-spread constraints and zonal persistent volumes can make an otherwise sufficient cluster impossible to scale in the required failure domain. A node group spanning several zones may not tell the autoscaler which zone a particular Pod needs, depending on the provider integration and group model. Where workloads have strict zone requirements, maintain compatible node groups per zone and validate the autoscaler’s balancing behavior instead of expecting a generic group to choose correctly.
Overprovisioning with low-priority placeholder Pods can reserve headroom and cause those placeholders to become unschedulable when real workloads arrive. If used, design it with explicit PriorityClasses, resource requests, and a budget: the placeholders consume real resources and must not be confused with production work. Cluster Autoscaler and HPA/queue-based workload scaling remain distinct controllers; observe each controller’s decisions in sequence.
Diagnose “no scale-up” and “no scale-down” systematically
For a pending Pod with no new node, check:
- Does the scheduler actually mark it unschedulable, and what are the exact reasons?
- Is the Pod eligible to trigger Cluster Autoscaler, considering its priority and configuration?
- Is there a configured node group that could satisfy all its resource, label, taint, topology, architecture, and volume constraints?
- Is that group below its maximum and within cloud quota?
- Did the cloud provider accept the resize, and did the machine boot and register?
- Did the new Node receive the labels, taints, devices, and allocatable resources modeled by the node-group template?
For an underutilized node that remains, check:
- Is the node in a group above its minimum and not protected by an exclusion annotation?
- Has it remained unneeded for the configured delay, and has a recent scale-up or failure introduced a cooldown?
- Can every movable Pod be placed elsewhere without violating resource, affinity, topology, or volume constraints?
- Does each relevant PDB permit the required eviction now, given healthy replicas and disruption allowance?
- Are DaemonSet, mirror, local-storage, or system Pods affecting utilization or eviction according to this CA version’s configuration?
- Did drain or provider termination fail, and what event or log records the failure?
Use the Cluster Autoscaler FAQ and the provider-specific documentation for exact flags, defaults, API permissions, node group labels, and version behavior. Heuristics and defaults can evolve, and managed Kubernetes distributions may package or configure the component differently. Do not transplant a flag from a different release or provider without checking the corresponding release branch and support matrix.
Version, permissions, and rollout discipline
The Cluster Autoscaler project recommends the latest release corresponding to the Kubernetes minor version running in the cluster. Its README warns that cross-version combinations are not generally tested in all environments. Check the upstream compatibility table at the time of upgrade and the cloud provider’s support policy. Upgrade the autoscaler as a control-plane component through the cluster’s supported mechanism, and canary the change where the platform allows it.
The autoscaler needs Kubernetes permissions to read workload and node state, and cloud permissions to inspect and resize configured groups. Grant only the actions and group discovery scope required for its provider integration. Review chart values, image tag, service account, cloud identity, and discovery filters together; a correctly running Pod with permissions to resize the wrong groups is still an unsafe deployment.
Before changing thresholds or group limits, capture status, events, desired group sizes, pending Pod reasons, and the current autoscaler image/version. Test a scale-up and a scale-down in a non-critical group, including a PDB-blocked case and a node that cannot be removed. Define how to restore the prior configuration and ensure another autoscaler does not race the rollback.
Cluster Autoscaler is functioning well when a truly schedulable pending workload prompts an appropriate configured capacity request, new nodes become usable within the service’s recovery objective, and unneeded nodes are removed only after their workloads can safely move. Judge it by these end-to-end outcomes, not by a single CPU chart or by the existence of an autoscaler Pod.
Related:
- Karpenter Node Lifecycle: NodePools, NodeClaims, Scheduling, Consolidation, and Disruption
- Kubernetes HPA Control Behavior: Metrics, Stabilization, and Scale Policies
Sources:
- Kubernetes node autoscaling: Cluster Autoscaler and node-group model
- Kubernetes Cluster Autoscaler README: operating model and version guidance
- Kubernetes Cluster Autoscaler FAQ: scale-up, scale-down, PDBs, and diagnosis
- Kubernetes Pod Disruptions: voluntary disruptions and disruption budgets
- Kubernetes Scheduling: resource requests and Pod placement