Kubernetes In-Place Pod Resize: CPU, Memory, Restart Policy, and Pending Changes
Safely resize running Kubernetes Pods with the /resize subresource, understand kubelet allocation states, restart policies, QoS limits, and rollout trade-offs.
Kubernetes can change a running Pod’s CPU and memory allocation without replacing the Pod. In-place Pod vertical scaling is useful when a workload needs a measured adjustment but a rollout would be unnecessarily disruptive. It is not a guarantee of zero impact: the kubelet must be able to allocate and apply the new values on the node, applications may need a restart to consume them, and Pod resource changes still affect scheduling capacity and namespace policy.
The key operational distinction is between the desired values in spec and the resources that the node has actually applied. A successful API request only records the desired resize. Operators must also check Pod status and container status before concluding that the change took effect. This article focuses on container-level CPU and memory resizing, which is stable starting with Kubernetes v1.35. Pod-level aggregate resource resizing is a separate capability that current Kubernetes documentation marks beta since v1.36; verify the exact release and feature support of your cluster before relying on it. See the Kubernetes Pod lifecycle and resize overview and the detailed container resource resize procedure.
What in-place resize does, and what it does not do
An in-place resize updates a running Pod through its /resize subresource. The API server validates the proposed resource values; the kubelet on the assigned node then attempts to allocate and apply the new CPU or memory configuration. A resize does not move the Pod to another node, add capacity to the cluster, or resize arbitrary resources such as GPUs, ephemeral storage, or a PersistentVolume. It also does not guarantee that the application dynamically changes its own worker count, heap sizing, cache limits, or other process-level settings.
The Pod’s QoS class is fixed at creation. A resize must remain compatible with that original class: a Guaranteed Pod must preserve matching CPU and memory requests and limits, a BestEffort Pod cannot acquire resource requests or limits, and a Burstable Pod cannot be changed into Guaranteed by making all requests equal all limits. An API request that violates the applicable invariant is rejected rather than silently changing the Pod’s QoS class.
In-place resize is available only for CPU and memory. Current Kubernetes documentation also lists important eligibility limits: Windows Pods, Pods managed by static CPU or memory manager policies, ephemeral containers, and non-restartable init containers cannot be resized in place. Existing resource requests or limits cannot simply be removed. Memory-limit reductions with no-restart policy are best-effort and can remain unapplied while current usage is too high; they are not a safe way to force memory usage down. Recheck the current documented limitations before treating any workload as eligible, since feature boundaries can evolve between releases.
Choose a restart policy per resource
Each container can declare a resizePolicy per resource. NotRequired is the default and asks Kubernetes to apply the change without restarting that container. RestartContainer allows Kubernetes to restart the affected container to apply the new value. If a request changes both CPU and memory, a restart-required policy for either changed resource can make the container restart. A Pod with overall restartPolicy: Never cannot specify a resize policy that requires a restart.
This is an application contract, not just a cluster tuning switch. CPU limits are generally enforced by the runtime without requiring the process to restart, but an application does not necessarily use newly available CPU to create more workers. Memory limits can be reflected in a container’s cgroup while a runtime or managed heap continues using its original startup configuration. If that process must reread memory sizing at startup, set a restart-required policy for memory and design the workload to tolerate the restart. Do not promise users that every resize is disruption-free merely because the Pod UID remains unchanged.
Here is a small example with an explicit policy. The CPU change is intended to apply in place; the memory change is allowed to restart the container. The requests equal limits for both resources, so the initial Pod is Guaranteed and a resize must preserve that relationship.
apiVersion: v1
kind: Pod
metadata:
name: resize-demo
namespace: operations
spec:
containers:
- name: app
image: registry.k8s.io/pause:3.8
resizePolicy:
- resourceName: cpu
restartPolicy: NotRequired
- resourceName: memory
restartPolicy: RestartContainer
resources:
requests:
cpu: 700m
memory: 200Mi
limits:
cpu: 700m
memory: 200Mi
Create the namespace and Pod through your usual reviewed deployment workflow. For a disposable test, for example:
kubectl create namespace operations
kubectl apply -n operations -f resize-demo.yaml
kubectl wait -n operations --for=condition=Ready pod/resize-demo --timeout=120s
Before changing anything, record the current desired and applied resources and restart count:
kubectl get pod resize-demo -n operations -o jsonpath='{.spec.containers[0].resources}'; printf '\n'
kubectl get pod resize-demo -n operations -o jsonpath='{.status.containerStatuses[0].resources}'; printf '\n'
kubectl get pod resize-demo -n operations -o jsonpath='{.status.containerStatuses[0].restartCount}'; printf '\n'
For Kubernetes clusters supporting the resize API and a kubectl client v1.32 or later, request a CPU change through the subresource:
kubectl patch pod resize-demo -n operations --subresource resize --type=merge \
-p '{"spec":{"containers":[{"name":"app","resources":{"requests":{"cpu":"800m"},"limits":{"cpu":"800m"}}}]}}'
Do not use a patch as a substitute for updating the owning Deployment, StatefulSet, or other workload template when the Pod is controller-managed. A direct Pod resize changes that Pod’s desired resources; it does not automatically update the template used for replacement Pods. If a later rollout creates a replacement from an unchanged template, the replacement can revert to the old allocation. Decide whether the change is an intentional per-Pod exception or the new durable workload configuration, and update the proper source of truth accordingly.
Verify the actual result, not just the accepted patch
Poll the Pod’s current specification, status, conditions, and events after requesting a resize:
kubectl get pod resize-demo -n operations -o yaml
kubectl describe pod resize-demo -n operations
kubectl get events -n operations --sort-by=.lastTimestamp
The spec shows desired resources. The container status reports resources applied by the node. Those values can differ while a resize is pending or in progress. Kubernetes exposes PodResizePending when the kubelet cannot immediately grant the change, and PodResizeInProgress while the application of a change is still underway. For pending state, inspect the reason and message: Deferred means the resize may become feasible later and kubelet will retry; Infeasible means the requested allocation cannot be satisfied on that node as currently configured. A successful API patch is therefore not proof that the cgroup has changed.
Where supported, compare Pod metadata.generation with status.observedGeneration to understand which desired Pod spec the kubelet has acknowledged. The observedGeneration recorded on resize conditions is useful when repeated changes arrive before an earlier resize completes. Prefer observing the fields from the live API instead of scripting against a remembered sample; exact status fields depend on Kubernetes version.
A Deferred resize is not an invitation to repeatedly patch increasingly large values. First establish whether the node has enough allocatable headroom, what other Pods reserve, whether cluster autoscaling can add capacity, and whether the new requests are justified by measurements. Increasing a request can make future placements harder even when the current Pod is already running. For workloads under namespace controls, confirm that ResourceQuota and LimitRange admit the new values. Track the pending condition’s age and alert when it outlasts the expected operational window.
If the desired change is infeasible, choose deliberately among restoring a feasible resource specification, freeing capacity, adding nodes through the normal capacity process, or replacing the Pod on a node with adequate capacity. Do not assume that eviction or rescheduling automatically occurs for an in-place resize. Kubernetes’ resize troubleshooting procedure shows the important failure signal: spec contains the new desired value while applied container status still contains the prior value.
Decide between an in-place change and a rollout
Use a Deployment or StatefulSet rollout when the change is part of the durable workload configuration, when other Pod fields must change at the same time, when the target environment does not support in-place resize, or when the application must be restarted in a controlled rollout. The workload controller’s strategy and readiness behavior make replacement explicit; a PodDisruptionBudget can help limit voluntary disruption but does not create spare capacity or guarantee that an update completes.
Use /resize for a bounded adjustment to an existing Pod when keeping that Pod identity is operationally valuable, the application can tolerate the selected restart behavior, and the node can satisfy the new resources. For a single-replica service, account for application behavior and its own readiness before selecting RestartContainer. A restart-required resource change may keep the Pod object but still create a period in which its container is restarting or not ready.
Vertical Pod Autoscaler and manual resize also need a clear ownership model. If an autoscaler, controller, GitOps reconciler, and operator can all write resource values, they may continuously overwrite one another. Choose one authoritative desired-state owner, configure its supported update mode, and test how it handles a resize that is deferred or requires a restart. The Vertical Pod Autoscaler behavior guide explains its recommendation and update modes; verify the installed VPA version’s compatibility and behavior rather than assuming that all modes use the Pod resize subresource.
Production rollout and rollback checklist
Before enabling in-place resize in an operational runbook:
- Confirm control-plane, kubelet, runtime, and
kubectlversions and verify the feature is supported on the actual node OS and resource-manager policies. - Exercise a small resize in a non-production namespace with the same QoS class, admission controls, and runtime configuration as the target workload.
- Record both desired and applied resource fields, container restart count, Pod conditions, and relevant events before and after the request.
- Test an increase that exceeds available node capacity so responders recognize
DeferredorInfeasiblebehavior and know the recovery path. - Test CPU-only, memory-only, and combined changes against the declared
resizePolicy; verify whether each case restarts the container. - Check quota, LimitRange, node allocatable resources, Pod placement, autoscaler behavior, and the workload controller’s source-of-truth template.
- Decide who may issue a resize, how GitOps or autoscaling controllers reconcile it, and how to return both the Pod and its template to a known-good state.
The rollback is another desired-state update, not an automatic undo. Restore the previous feasible requests and limits through the /resize subresource, then verify that status reflects them. If the resize is infeasible because of capacity or policy, first resolve that constraint or use the workload controller to replace the Pod in a deliberately planned rollout. Keep a record of the requested generation, the admitted spec, the applied container status, and the final recovery action.
In-place Pod resize is a useful control for carefully managed resource changes, but it is not equivalent to a node upgrade, automatic vertical autoscaling, or a zero-downtime rollout. Production safety comes from understanding the two resource states, respecting the Pod’s original QoS and node constraints, choosing a restart policy the application can handle, and verifying the actual result before declaring the change complete.
Related:
- Kubernetes Vertical Pod Autoscaler: Recommendation, Update Modes, and Safe Rollouts
- Kubernetes Pod Priority and Preemption: Scheduling Order, Victims, and PDB Limits
Sources: