Kubernetes Vertical Pod Autoscaler: Recommendation, Update Modes, and Safe Rollouts
Operate Kubernetes VPA safely by separating recommendations from mutations, choosing update modes, bounding resources, and planning disruption and HPA interaction.
The Kubernetes Vertical Pod Autoscaler (VPA) changes the resource requests, and optionally limits, of Pods based on observed workload behavior. It does not add replicas, create cluster capacity, or make an unsafe eviction safe. Production use requires separating three questions: what resources the recommender suggests, when the admission controller applies a suggestion to a new Pod, and whether the updater is allowed to disrupt an existing Pod to change its resources.
VPA is an add-on project, not a built-in Kubernetes control-plane feature. The cluster must have a compatible VPA installation, its custom resource definition, and the components required by the selected update mode. Check the release and installation instructions for the VPA version actually running; upstream behavior and in-place update capabilities evolve independently from Kubernetes.
The components have separate responsibilities
The recommender observes resource usage and publishes recommendations in a VerticalPodAutoscaler object’s status. The updater evaluates whether existing Pods should be changed and, in restart-based modes, requests eviction. The admission controller applies recommendations to Pods as they are created. A VPA object connects these components to a target workload and optional resource policies.
The VPA status is a recommendation, not proof that a new request will fit on a node. A recommended request may be larger than current free allocatable capacity, blocked by a namespace quota, or incompatible with a storage or topology constraint. Admission can produce a Pod with new requests that remains Pending. Treat recommendation review, Pod admission, and scheduling as distinct checkpoints.
Recommendations include target, lower-bound, and upper-bound values for controlled resources. The target is the current recommended amount; bounds communicate the recommender’s range, not a guarantee of safe application behavior. Resource policy minAllowed and maxAllowed constrain recommendations. They should encode reviewed operational boundaries rather than conceal missing capacity or poor telemetry.
Start with observation and explicit bounds
The example deliberately uses Off mode. It allows the team to inspect recommendations without automatically modifying Pod resources. Replace the namespace, workload name, and example bounds after measuring the service; the quantities below are schema examples, not universal sizing guidance.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: checkout-api
namespace: production
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: checkout-api
updatePolicy:
updateMode: Off
resourcePolicy:
containerPolicies:
- containerName: "*"
controlledResources: ["cpu", "memory"]
controlledValues: RequestsOnly
minAllowed:
cpu: 100m
memory: 128Mi
maxAllowed:
cpu: "2"
memory: 2Gi
Using RequestsOnly makes the policy’s ownership clear when the application team intentionally manages limits separately. The API also supports RequestsAndLimits, which is the default if controlledValues is omitted. VPA resource policies can opt a container out or bound the resources it controls; review sidecars as carefully as the main process, because each container has its own resource requirements.
After applying a VPA, inspect both spec and status and compare recommendations with monitoring data from representative traffic:
kubectl get vpa checkout-api -n production -o yaml
kubectl describe vpa checkout-api -n production
kubectl get pods -n production -l app=checkout-api -o wide
kubectl top pods -n production -l app=checkout-api
Metrics from a short or unrepresentative period can miss infrequent jobs, startup peaks, seasonal load, batch windows, and memory growth. Decide how much history the installed recommender can observe and what periods the application must survive. Validate recommendations against application-level latency, queue depth, throttling, out-of-memory kills, and known workload cycles rather than copying a single number from a dashboard.
Select an update mode as a disruption policy
Specify updateMode explicitly. Upstream API documentation currently lists these modes, but feature availability and defaults must be verified against the deployed VPA release:
- Off calculates and exposes recommendations without automatically changing Pod resources. Use it for an observation phase and as a safe rollback mode.
- Initial applies recommendations when a Pod is created, but does not update that Pod later. It is useful when a separate rollout process controls replacement timing.
- Recreate applies recommendations to new Pods and may evict existing Pods so that replacements receive updated resources. It can cause real service disruption and may wait on disruption constraints.
- InPlaceOrRecreate attempts to resize existing Pods in place where the required Kubernetes and VPA capabilities are enabled, then can fall back to eviction and recreation. It is not a promise of zero disruption.
- InPlace is an eviction-free update mode in current upstream VPA documentation, but is feature-gated and marked alpha there. An infeasible resize is deferred and retried; it is not a guarantee that the desired resources will eventually fit.
Auto is deprecated in current upstream VPA documentation and is described as equivalent to Recreate there. Prefer an explicit mode so a future release cannot silently change the operational intent. Also inspect the VPA release notes and installed CRD before applying a mode that may not be recognized by a managed distribution.
For restart-based updates, an updater uses Pod eviction so that the workload controller can recreate the Pod and the admission controller can apply the revised resources. VPA’s upstream quick start says the Recreate mode respects a PodDisruptionBudget when one is defined. That still does not guarantee timely replacement: a restrictive budget, too few healthy replicas, quota exhaustion, or no schedulable capacity can stall the update. A PDB expresses voluntary disruption allowance; it does not manufacture a replacement Pod or a free node.
Roll out changes in controlled stages
- Install and validate VPA components in a non-production or dedicated canary cluster. Confirm that the CRD, recommender, updater, and admission controller are all healthy, and that their logs and metrics are available.
- Create a VPA in Off mode for a representative workload. Observe recommendations across normal traffic, peak traffic, restart, and batch windows. Review every controlled container and compare the values with requests, limits, quotas, and node allocatable capacity.
- Bound the policy intentionally. Too-low maxAllowed can leave a memory-hungry process vulnerable to OOM termination; too-high maxAllowed can make replicas unschedulable. minAllowed can prevent the recommender from following a transient low sample into an unsafe configuration, but it must be justified by evidence.
- Choose Initial when admission-time mutation is acceptable and another release mechanism will replace Pods. Choose Recreate only after rehearsing evictions, ensuring adequate healthy replicas, and validating budget and capacity behavior. Prefer an explicitly supported in-place mode only after checking feature gates and resize status semantics in the exact VPA/Kubernetes versions.
- Change one workload or canary slice at a time. Observe Pod resource specs after admission, Ready replica count, HPA status, scheduling events, restart reasons, throttling, memory pressure, and application SLOs. Keep a rollback that can stop further VPA updates and restore reviewed workload-template values.
VPA changes actual Pod resource fields; it does not rewrite the Deployment’s Pod template. This separation matters for GitOps and rollbacks. The Deployment may continue to declare its original requests while VPA-mutated Pods carry new values. Document which controller owns each field, ensure drift tooling will not fight the VPA-managed Pod state, and verify what resources newly created Pods receive after any template update.
Coordinate carefully with HPA and resource limits
VPA does not scale replica count. HPA does. For CPU utilization targets, HPA evaluates usage relative to requested CPU; if VPA changes that request, the utilization percentage can change even when application CPU consumption does not. Two control loops may then react to each other’s changes. The VPA project FAQ documents a memory-only VPA with CPU-based HPA as one way to separate the controlled resources. Other metric designs are possible, but must be load-tested for feedback and scale stability.
VPA can manage CPU, memory, or both. Decide whether it owns requests only or both requests and limits. Limits may affect CPU throttling and memory OOM behavior; changing a memory limit downward while a process is using more memory can have immediate consequences. Preserve intentional request-to-limit relationships and test the resulting QoS class. Do not assume a recommendation is safe merely because it lies inside a syntactic min/max range.
If a namespace uses ResourceQuota or LimitRange, inspect the post-admission Pod spec and quota usage after VPA changes. A VPA recommendation does not bypass API admission. A new resource request may be rejected, may fit the quota but not any node, or may cause a rollout to remain below its desired availability. Diagnose admission, quota, scheduling, and application health separately.
Troubleshoot by stage
- Missing recommendation: check VPA status conditions, recommender logs, the target reference, and whether the target owns the Pods directly in a way VPA supports.
- Recommendation changes but Pod resources do not: confirm updateMode, updater health, admission-controller webhook health, and the VPA container policy. Off and Initial are not running-Pod resize modes.
- Recreate updates stall: inspect PodDisruptionBudget status, updater events, healthy replicas, and available node capacity. Avoid weakening the budget until you know the application can tolerate the interruption.
- New Pods remain Pending: compare their admitted requests with allocatable capacity, quota, affinity, topology constraints, and volume placement. A larger VPA request may be the direct cause.
- Pods restart or fail after a change: correlate the new memory limit and request with OOM and restart evidence, and correlate CPU request changes with throttling and HPA metrics.
- In-place changes defer or fail: inspect Pod resize conditions and messages, feature-gate configuration, node headroom, resource resize policy, and the exact fallback behavior documented by the installed release.
Keep recommendation history, VPA status, updater events, admitted Pod specs, and application metrics together in change reviews. Alert on stalled recommendation freshness, repeated evictions, unresolved resize errors, Pending replacements, OOM increases, and HPA behavior changes. The production acceptance criterion is not that the VPA object exists; it is that the recommendation is explainable, each mutation is bounded, replacements remain schedulable, and application SLOs remain inside their error budget.
Related:
- Kubernetes HPA Control Behavior: Metrics, Stabilization, and Scale Policies
- Kubernetes QoS Classes: How Requests and Limits Shape Eviction and Scheduling
Sources: