Kubernetes HPA Control Behavior: Metrics, Stabilization, and Scale Policies
Understand how Kubernetes HPA turns metrics into replica recommendations, handles incomplete data, limits scaling velocity, and avoids scale-down flapping.
A HorizontalPodAutoscaler is a sampled control loop, not a continuously reacting load balancer. It reads metrics, calculates a replica recommendation, applies safety rules, and writes the result through the target workload’s scale subresource. A successful HPA therefore proves only that Kubernetes can request a replica count. It does not prove that Pods can schedule, become Ready, or relieve the application bottleneck.
This article focuses on that control path and its failure modes. For initial installation and a basic manifest, see the HPA setup guide.
Inputs: a scale target and a working metrics API
Use the stable autoscaling/v2 API and a target that implements the Kubernetes scale subresource, such as a Deployment or StatefulSet. The HPA uses the target’s Pod selector to identify the Pods whose metrics it will evaluate.
Resource metrics such as CPU and memory come through metrics.k8s.io, commonly provided by Metrics Server, which is a separate cluster add-on and is not guaranteed to be installed. Check the API version as well as the API group: the Metrics API graduated to stable metrics.k8s.io/v1 in Kubernetes v1.37, but that release’s HPA controller still consumes metrics.k8s.io/v1beta1; kubectl top supports both versions and prefers v1 when available. Custom metrics and external metrics require their corresponding aggregated APIs and adapters. The aggregation layer and APIService registration must be healthy; a running metrics-server Deployment alone does not prove that the HPA can read fresh data. Verify the version your control-plane release consumes, the APIService health, and an end-to-end HPA metric value.
For CPU utilization targets, utilization is measured relative to resource requests. If a relevant container in a selected Pod lacks a CPU request, the Pod’s CPU utilization is undefined for that metric and the HPA cannot make the expected calculation. Include injected sidecars in the review: their CPU use and requests contribute to Pod-level utilization. Since Kubernetes v1.30, a stable ContainerResource metric can target the application container instead when Pod-level utilization hides its demand. Keep the container name accurate across rollouts; during a rename, configure the HPA to track old and new names until the rollout completes.
The control loop runs periodically in kube-controller-manager; the documented default sync period is 15 seconds and the operator can change it. Metrics collection, API aggregation, sample windows, and Pod startup add latency beyond that interval.
From a metric to a replica recommendation
For a per-Pod metric with complete data, the simplified calculation is:
desiredReplicas = ceil(currentReplicas * currentMetric / targetMetric)
With six replicas averaging 78% CPU utilization against a 65% target, the raw recommendation is ceil(6 * 78 / 65), or eight. This is the starting recommendation, not necessarily the final scale: HPA also applies tolerance, min/max bounds, stabilization, and scaling-rate policies. For utilization, the average CPU value is expressed as a percentage of requested CPU; raw metrics use their configured units.
With multiple metrics, HPA calculates a recommendation for each and selects the largest. This protects scale-up when one signal sees more demand than another. Conversely, if a metric cannot be fetched and the available metrics recommend scaling down, HPA skips that scale-down because it cannot safely conclude that demand is low. A metrics error therefore does not always mean “no scaling”; the direction and the other metrics matter.
The controller is deliberately conservative around incomplete Pod data. Missing metrics are treated as consuming the target amount during a scale-down and as zero during a scale-up, reducing the size of the change. For CPU, initializing or not-yet-ready Pods can also be set aside; if a scale-up is otherwise indicated, their assumed zero usage dampens the recommendation. The documented controller defaults for CPU startup handling are a 30-second initial readiness delay and a five-minute CPU initialization period. These are cluster-wide controller-manager settings, not per-HPA fields. A truthful startup/readiness design matters because a premature Ready signal can make startup CPU samples look like steady-state demand.
The default tolerance is 10% around the target, so small metric swings do not trigger replica churn. Per-direction behavior.scaleUp.tolerance and behavior.scaleDown.tolerance are stable in Kubernetes v1.37; the feature first appeared in v1.33 behind HPAConfigurableTolerance. On older clusters, check the version, feature-gate state, and API schema before using those fields. When they are unset, the cluster-wide default can be changed through the kube-controller-manager tolerance flag.
Stabilization and rate policies solve different problems
A stabilization window smooths recommendations; a scaling policy limits how quickly replicas may change. They are not interchangeable. For scale-down, the default stabilization window is 300 seconds. HPA considers recent recommendations and uses the highest recommendation in that rolling interval, avoiding a drop immediately after a short low-metric sample. This is not a simple five-minute timer that starts once and then guarantees a downscale. The current metric must still recommend fewer Pods, and minReplicas and rate policies still apply.
Policies of type Pods cap an absolute replica change; Percent caps a relative change. periodSeconds describes the lookback period over which permitted changes are evaluated, not the HPA reconciliation interval. With multiple policies, selectPolicy: Max (the default) allows the largest change among them; Min chooses the more restrictive change, and Disabled turns off scaling in that direction.
In Kubernetes v1.37, the documented default scale-up behavior has no stabilization window and allows either four Pods or 100% growth per 15-second policy period, selecting the larger change. Default scale-down allows up to 100% per 15 seconds, with a 300-second stabilization window. Defaults are not a capacity plan. Slow startup, expensive connection pools, or a fragile downstream service can justify a slower custom scale-down rate and a deliberately bounded scale-up rate.
This example assumes a checkout Deployment already exists in production, its containers declare CPU requests, and the cluster exposes working resource metrics. Its scale-down policy is intentionally slower than Kubernetes’ default:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: checkout
namespace: production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: checkout
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
behavior:
scaleUp:
stabilizationWindowSeconds: 0
selectPolicy: Max
policies:
- type: Percent
value: 100
periodSeconds: 60
- type: Pods
value: 4
periodSeconds: 60
scaleDown:
stabilizationWindowSeconds: 300
selectPolicy: Min
policies:
- type: Percent
value: 10
periodSeconds: 60
- type: Pods
value: 2
periodSeconds: 60
At 10 replicas, these scale-down rules allow the smaller change from one replica (10%) or two Pods. Review those values against workload warm-up, dependency limits, and the recovery objective; do not copy them as universal production defaults.
Diagnose the recommendation separately from execution
Start with the HPA’s reported metric, target, current and desired replica counts, conditions, and Events:
kubectl get hpa checkout -n production -w
kubectl describe hpa checkout -n production
kubectl get hpa checkout -n production -o yaml
kubectl top pods -n production --containers
kubectl get apiservice | grep 'metrics.k8s.io'
If the target is <unknown>, trace the relevant aggregated API and adapter, then verify that the selector matches the intended Pods and that every required CPU request exists. If metrics are valid but desiredReplicas does not rise, compare the measured value with the target and tolerance, check readiness handling, and confirm that maxReplicas or a scale policy is not limiting the recommendation. If HPA desires more replicas but they stay Pending or unavailable, investigate Deployment events, scheduling constraints, quota, image pulls, node capacity, and readiness; HPA cannot provision capacity by itself.
For oscillation, compare fresh metric timestamps and observed replicas with the tolerance, recommendation window, policy history, and actual Pod-ready time. An HPA status value is a recent controller observation, not necessarily the same instant as a dashboard’s metric sample.
A useful acceptance test is repeatable and checks both control and service outcomes:
- Under bounded load, record metric freshness and confirm the reported current value exceeds the target by more than tolerance. Verify that
desiredReplicasfollows the expected ratio, subject to configured bounds and policies. - Confirm the requested replicas actually become Ready and that the service’s latency or queue-age objective improves. A replica increase without this result is not successful autoscaling.
- Remove load and verify that scale-down respects the configured rate and never drops below the highest recommendation retained by the stabilization window.
- In a staging test, make the metric source unavailable and drive demand in both directions. Confirm that a failed metric is visible, that an invalid metric does not cause an unsafe downscale, and that alerts distinguish a max-replica ceiling from a metrics failure.
Treat HPA as one component of a feedback system. Choose metrics that respond predictably to adding replicas, leave enough cluster and dependency capacity for the configured maximum, and measure time-to-recommendation separately from time-to-Ready. For event-driven activation or scale-to-zero, the HPA has a different operating contract; the KEDA deep dive covers that boundary.
Related:
- How to Set Up Horizontal Pod Autoscaling in Kubernetes
- KEDA Event-Driven Autoscaling: ScaledObjects, ScaledJobs, HPA, and Safe Scale-to-Zero
Sources:
- Horizontal Pod Autoscaling documentation
- Kubernetes v1.33: HorizontalPodAutoscaler Configurable Tolerance
- HorizontalPodAutoscaler v2 API reference
- Kubernetes observability: metrics API versions and HPA compatibility
- Kubernetes v1.37: Metrics API graduates to stable
- Resource metrics pipeline
- Kubernetes v1.37 replica calculator source
- Kubernetes v1.37 resource-metrics client source
- Kubernetes v1.37 HPA defaulting source