Kubernetes Pod IP Capacity: Plan CIDRs and Diagnose CNI Exhaustion
Plan Kubernetes Pod IP capacity across Pod, node, and Service ranges; diagnose CNI allocation failures, subnet exhaustion, warm pools, and safe remediation.
Kubernetes can report healthy nodes and available CPU while new Pods remain stuck because the networking layer cannot allocate an address. The phrase “IP exhaustion” is ambiguous: a cluster may run out of Pod addresses, node addresses, cloud network-interface slots, per-node Pod capacity, or Service addresses. Each pool is allocated by a different component and has a different safe recovery path.
This guide treats address capacity as an end-to-end system: Kubernetes Pod CIDRs, node and Service ranges, the Container Network Interface (CNI) plugin, provider-specific IPAM, per-node Pod limits, and warm address pools. The exact allocation model is CNI- and cloud-specific. Use the provider’s current capacity tools and node configuration as the source of truth; formulas below are planning prompts, not universal provider limits.
Separate the address pools before changing anything
The Kubernetes network model gives each Pod a unique cluster-wide address, but Kubernetes does not itself reserve one universal address block in every cluster. The network plugin and infrastructure determine how Pod addresses are allocated and routed. A useful incident inventory separates these resources:
| Capacity pool | What consumes it | Common owner |
|---|---|---|
| Pod addresses | Pod network interfaces and, depending on the CNI, preallocated warm addresses or prefixes | CNI/IPAM and cloud networking integration |
| Node addresses | Worker node primary interfaces and secondary interfaces | Cloud subnet, on-premises IPAM, or node network configuration |
| Service addresses | ClusterIP allocations from the configured Service CIDR | Kubernetes API server Service IP allocator |
| Per-node Pod slots | Regular Pods plus system and DaemonSet Pods counted by the kubelet’s configured limit | Kubelet and cluster provisioning system |
| Interface and prefix capacity | ENIs, addresses per interface, delegated prefixes, routes, or eBPF/IPAM entries | Instance type, cloud provider, or CNI implementation |
The pools can be coupled but should not be conflated. A provider may allocate each Pod IP from a cloud subnet; another CNI may allocate a per-node CIDR from a cluster range and route it across nodes. Some CNIs reserve addresses for warm capacity. A node may have plenty of Pod CIDR space but hit its interface attachment limit, or have unused instance slots but no free addresses in the subnet. Dual-stack clusters need capacity planning independently for IPv4 and IPv6.
Keep Pod, node, Service, VPC/VNet, on-premises, VPN, and peered-network ranges non-overlapping where they must be routed together. An overlap can cause a route ambiguity even when every pool appears numerically large enough. Provider-reserved addresses, subnet masks, and prefix-allocation rules also change usable capacity; do not estimate usable addresses by subtracting a universal constant from every CIDR.
Plan for peak nodes and transient overlap
Start with maximum, not average, concurrency. For each node pool, record the minimum and maximum node count, per-node Pod limit, DaemonSet and system Pod count, autoscaler behavior, upgrade surge, disruption policy, and whether replacement nodes can coexist temporarily with old nodes. Then estimate the Pod-side capacity that must be available at peak:
peak Pod slots by pool =
peak simultaneous nodes (including upgrade / scale-out overlap)
* configured Pods-per-node ceiling
workload slots =
peak Pod slots
- expected system and DaemonSet Pods
- explicit safety headroom
The first number is a conservative slot bound, not always the exact number of IPs the provider will allocate. For a CIDR-per-node design, the allocator may reserve an entire block for each node, so unused addresses inside a node block are not available to other nodes. For an ENI/IPAM design, warm addresses may consume subnet capacity before a Pod is created. Translate slot requirements through the actual plugin’s allocation model.
Include upgrade and failure scenarios. A cluster that can run 500 Pods after it has converged may still need extra addresses while a node is drained and a replacement is provisioning. Autoscaling can create several nodes concurrently; a burst of Jobs can consume warm pools and trigger cloud API throttling. Capacity reviews should model the largest supported node shape and the least generous network limit across a heterogeneous node group.
Use distinct budgets for Pod addresses, node addresses, and Service ClusterIPs. The Service CIDR usually grows independently of worker subnet use. Node IP exhaustion can block node provisioning even if there are spare Pod addresses. Pod range exhaustion may prevent sandbox creation for already scheduled Pods. Keep a documented reserve and alert before a pool reaches its hard boundary, because expanding some ranges requires control-plane or node-pool recreation.
Diagnose where allocation failed
First determine whether the scheduler could not place the Pod or whether networking failed after placement. A scheduler capacity limit commonly leaves a Pod Pending with an event such as Too many pods. In contrast, a Pod may already be bound to a node and remain in ContainerCreating while kubelet and the runtime ask the CNI to create its sandbox. Events may include FailedCreatePodSandBox or an IPAM-specific error. Exact messages depend on the runtime and plugin.
kubectl get pods -A -o wide
kubectl describe pod -n production api-7f8c9d7d8b-x4p2q
kubectl describe node worker-pool-a-3
kubectl get events -A --sort-by=.lastTimestamp
Check the Pod’s node assignment, PodScheduled condition, recent sandbox events, and the node’s .status.allocatable.pods. If the scheduler reports no feasible node, investigate maxPods, requests, taints, affinity, and capacity before blaming IPAM. If a bound Pod fails CNI setup, examine the node networking DaemonSet logs, CNI/IPAM metrics, cloud API errors, interface address counts, prefix availability, route limits, and free addresses in the exact subnet or range associated with that node.
For a provider-managed CNI, use its supported diagnostics rather than manually editing plugin state files. Correlate timestamps between kubelet/runtime events, CNI logs, cloud audit records, and subnet metrics. A transient cloud API throttling error is not the same as a depleted CIDR; both can delay address allocation, but the fixes differ. Avoid deleting CNI Pods or manually freeing addresses until you understand which component owns the allocation and whether it can safely reconcile that change.
Understand allocation mode and warm pools
Address allocation strategy affects both launch latency and subnet consumption. A CNI may allocate one secondary address when a Pod starts, attach an interface with multiple addresses, prefetch a warm pool, or delegate a prefix. A warm pool can reduce the time between scheduling and network readiness because the CNI has addresses ready locally, but those unassigned addresses still consume the underlying network range.
On Amazon EKS, the Amazon VPC CNI configuration exposes targets such as WARM_IP_TARGET, MINIMUM_IP_TARGET, and WARM_PREFIX_TARGET; the effect depends on secondary-IP, prefix-delegation, custom-networking, and other configuration. Prefix delegation can improve allocation efficiency, but each required contiguous prefix must be available in the subnet. Existing fragmented address use can prevent prefix assignment even when the arithmetic count of free addresses looks adequate. Check the CNI version, instance network limits, subnet free space, and provider guidance before enabling it.
GKE VPC-native clusters use subnet secondary ranges for Pod addresses and separate ranges for node and Service addressing. Exhausting the Pod secondary range is different from exhausting the primary node range. Google documents options such as adding Pod ranges in supported configurations or recreating a cluster with a larger range; some remedies require node pool recreation or other planned changes. Do not copy a GKE range formula into a different networking mode or another provider.
For every CNI, inspect the actual configuration in the running DaemonSet or managed node image and compare it with the infrastructure source of truth. A warm-pool target that is too high can strand addresses in advance; one that is too low can increase Pod start latency and cloud API calls. Tuning is a throughput-versus-reserved-capacity trade-off. Load test bursts and watch both address inventory and sandbox latency before changing defaults.
Remediate with the right owner and blast radius
Choose remediation only after identifying the exhausted pool:
| Finding | Candidate response | Risk to check first |
|---|---|---|
| Scheduler reaches the kubelet’s Pod-slot limit | Scale the node pool, choose a supported larger node, or adjust the Pod limit through the platform workflow | CNI/interface capacity may not support the new limit; managed node groups may need replacement |
| Provider Pod subnet/range is exhausted | Expand or add a supported Pod range, migrate to a new range, or add a compatible node pool | Existing Pods, routes, firewall rules, peering, and upgrade overlap must all remain reachable |
| Warm addresses dominate available subnet space | Carefully reduce CNI warm targets or change allocation mode after testing | Cold Pod starts, cloud API throttling, and IP assignment latency may increase |
| Node subnet has no addresses for scale-out | Expand the node subnet or provision nodes in another supported subnet | Routing, availability-zone placement, load balancers, and control-plane access can change |
| Service range is near exhaustion | Follow the platform-specific Service CIDR expansion or migration procedure | ClusterIP addresses are embedded in Service state and dependent policy or external systems |
| CNI cloud calls are throttled but ranges remain | Address API quotas, permissions, retry behavior, or allocation bursts | Raising concurrency can worsen provider throttling and produce a retry storm |
Do not reduce maxPods on a live production fleet as an immediate fix without checking rollout semantics. Existing nodes may keep their old limit while replacement nodes receive the new one, and some provider products only apply the change after a node image or launch-template rollout. Likewise, a larger Pod range does not automatically update every node’s per-node block assignment or a CNI’s address pool. Use the provider’s documented procedure, reserve capacity for migration overlap, and test rollback before moving workloads.
If only a subset of workloads needs a new network model, a separate node pool or cluster can reduce migration risk. That isolation adds its own costs: duplicate system services, different capacity, routing and policy integration, and operational complexity. Treat it as an architecture decision, not an ad hoc response to an exhausted subnet.
Monitor leading indicators
Alert on remaining addresses in every relevant pool, not just total VPC free space. Track the distinction between allocated-to-Pod addresses, addresses held warm, addresses reserved for node interfaces, and free addresses that are actually usable under the plugin’s allocation mode. For prefix allocation, monitor free contiguous prefixes at the required size, not only aggregate free IPv4 count.
Useful operational indicators include:
- CNI IPAM allocation failures, retries, latency, cloud API throttling, and warm-pool size.
- Pod sandbox creation failures and time spent in
ContainerCreatingafter scheduling. - Node
.status.allocatable.pods, actual Pod count, and headroom by node pool. - Free addresses and network-interface capacity per subnet, zone, and node type.
- Pod, node, and Service CIDR utilization, including IPv4 and IPv6 separately.
- Upgrade surge, autoscaler max size, and the address reserve needed for simultaneous replacement nodes.
Test scale-out under burst load and simulate an unavailable subnet or zone. Verify not only that Pods eventually receive addresses, but also that DNS, service routing, network policies, load balancers, and external routes continue to work with the new address ranges. Update IPAM diagrams, runbooks, IaC validation, and alert thresholds whenever node shapes, CNI mode, or network allocation settings change.
Production acceptance checklist
- Pod, node, and Service address pools are documented separately, including provider reservations, routing overlap, and dual-stack capacity.
- Peak planning includes max nodes, Pod limits, DaemonSets, warm allocations, upgrade surge, autoscaler growth, and migration overlap.
- Runbooks distinguish scheduler
Too many podsfrom post-bind CNI sandbox/IPAM failures and name the owner of each pool. - CNI allocation mode, warm-pool behavior, interface/prefix limits, and cloud API quotas have been validated against the deployed plugin and node type.
- Range expansion and node-pool migration procedures are tested with rollback and enough capacity to keep services healthy.
- Monitoring covers remaining and warm addresses, prefix fragmentation, allocation errors, Pod startup latency, and separate node/Pod/Service ranges.
- Routing, DNS, NetworkPolicy, Service access, and external connectivity are tested after any range or IPAM change.
Kubernetes IP capacity is not a single CIDR-size calculation. Production safety comes from tracing each Pod address from the scheduler’s node limit through kubelet, CRI, CNI/IPAM, and the infrastructure pool that ultimately supplies it. Size and observe every layer, then remediate the layer that actually ran out.
Related:
- Understanding Kubernetes Networking: Services, kube-proxy, and CNI Plugins
- Fixing Pods Stuck in Pending State in Kubernetes
Sources:
- Kubernetes: The Kubernetes Network Model
- Kubernetes: Cluster Networking and IP Address Ranges
- Container Network Interface: Specification
- Amazon EKS: Amazon VPC CNI Best Practices
- Amazon EKS: Optimize IP Address Utilization
- Google Kubernetes Engine: VPC-Native Clusters and Alias IPs
- Google Kubernetes Engine: Troubleshoot IP Address Management