Kubernetes Pod Priority and Preemption: Scheduling Order, Victims, and PDB Limits
Design Kubernetes PriorityClasses and preemption policies with precise feasibility, termination, PDB, quota, and operational trade-offs.
Pod priority is an ordering and recovery policy. It tells the Kubernetes scheduler which pending work should be considered first and, when preemption is allowed, which lower-priority Pods may be removed to make a candidate node feasible. It is not a reservation, a placement guarantee, or a general override for admission and topology rules.
The operational cost is borne by the victims: their containers receive termination time, their controllers must create replacements, and those replacements still need quota and a schedulable node. A PriorityClass is therefore a cluster-wide contract about which work may displace which other work. Give it explicit owners and a tested recovery objective rather than using priority as a shortcut for missing capacity.
Define priority classes as cluster policy
A PriorityClass is a non-namespaced object that maps a human-readable name to an integer value. Higher values mean higher priority. Ordinary PriorityClass values can range from -2,147,483,648 through 1,000,000,000; larger values are reserved for built-in system-critical classes. Do not copy system-critical classes into application workloads.
The example below assigns an illustrative value, not a Kubernetes-recommended production number. Choose values from an organization-wide policy so that application teams cannot independently label every workload as critical. Set globalDefault deliberately: at most one class can be the global default, and it affects Pods created after it is set, not Pods that already exist. If no default class exists, Pods without priorityClassName receive priority zero.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: payments-critical
value: 100000
globalDefault: false
preemptionPolicy: PreemptLowerPriority
description: "Reserved for approved payment request-serving workloads."
Workload templates consume the class by name:
apiVersion: apps/v1
kind: Deployment
metadata:
name: payments-api
namespace: production
spec:
replicas: 3
selector:
matchLabels:
app: payments-api
template:
metadata:
labels:
app: payments-api
spec:
priorityClassName: payments-critical
containers:
- name: api
image: registry.example.net/payments-api:reviewed
resources:
requests:
cpu: 250m
memory: 512Mi
This changes scheduling priority but does not automatically raise CPU or memory requests. PriorityClass resolution occurs during Pod admission; if the named class does not exist, Kubernetes rejects the Pod. Manage class creation and workload authorization with RBAC, and consider ResourceQuota scoped to PriorityClass to limit how much high-priority capacity a namespace can consume.
Understand when preemption can help
The scheduler first tries to place a pending Pod on feasible nodes. If no node can fit it, preemption may be considered when removing one or more lower-priority Pods would make a node feasible. Victims must have lower priority than the preemptor. The scheduler can remove only a subset when that is enough; it does not need to evict every lower-priority Pod from the chosen node.
Preemption is useful for a resource shortage that victim removal can actually repair. If the Pod cannot tolerate a node taint, requires a missing topology label, conflicts with a hard node selector, or cannot attach its volume in a candidate domain, removing an unrelated workload does not make that node valid. Similarly, a namespace quota rejection happens at API admission, before scheduling; raising priority does not bypass quota, LimitRange, or admission policy.
Inter-Pod affinity creates an especially subtle boundary. A pending Pod cannot be scheduled on a candidate node if doing so depends on a lower-priority Pod that preemption would remove. The scheduler does not evict that Pod and then assume the affinity rule remains satisfied. It may look for another node, but there is no guarantee one exists.
Preemption is also not an instantaneous reservation. When the scheduler selects victims, the preemptor may be associated with a nominated node while those Pods terminate. During that interval, another pending Pod may be scheduled there, and the original preemptor must be evaluated again. Even after victim removal, competing work, topology changes, or insufficient remaining capacity can prevent the pending Pod from running.
Balance queue preference against eviction
The default preemption policy, PreemptLowerPriority, permits a class to evict lower-priority Pods when the scheduler needs to make room. A class with preemptionPolicy: Never is still prioritized in the scheduling queue but cannot itself preempt running Pods. It waits for enough resources to become free naturally and is subject to scheduler backoff; a higher-priority Pod may still preempt it.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: batch-important-nonpreempting
value: 1000
globalDefault: false
preemptionPolicy: Never
description: "Queue batch work ahead of lower priority without evicting running Pods."
Non-preempting priority is appropriate when order matters but discarding other work is not acceptable. It does not reserve future CPU or memory, and a high-priority Pod that cannot fit can still wait indefinitely if no capacity becomes available. Combine it with a credible capacity plan and queue-age alert.
Avoid creating a very high priority class for every service. If many Pods share the same top value, they cannot preempt one another, so priority alone does not resolve contention among them. A broken or compromised tenant with permission to request an extreme class can starve ordinary workloads. Restrict class creation and use, use class-scoped quotas where appropriate, and review the exact Pod templates that inherit a class.
Treat victim termination and PDBs realistically
Preempted Pods receive their graceful termination period before they are forcibly terminated. The scheduler can continue processing other pending Pods while it waits. The remaining grace period therefore contributes to recovery latency, and a preemptor may not run immediately after the decision to preempt. Set termination behavior according to actual data-safety and request-draining requirements; setting a near-zero grace period simply to accelerate preemption can truncate work and connection draining.
Kubernetes considers PodDisruptionBudgets during preemption, but respecting them is best effort. The scheduler tries to find a set of victims that does not violate applicable PDBs. If it cannot find such victims and preemption is needed, it may still preempt Pods in violation of those budgets. A PDB is not an absolute shield against scheduler preemption.
This differs from a planned node drain that uses the eviction API, where a restrictive PDB can refuse an eviction and stall maintenance. Do not infer from a healthy PDB status that a lower-priority workload is protected from every preemption scenario. Protect critical services with redundant replicas, separation across failure domains, adequate headroom, and tested recovery, not just a PDB.
Priority also interacts with node-pressure eviction but is not the only decision signal. The kubelet considers whether a Pod’s usage exceeds its requests, priority, and usage relative to requests when ranking Pods for node-pressure eviction. Scheduler preemption and kubelet pressure eviction are distinct mechanisms with different triggers. A high PriorityClass is not a blanket immunity from a node failing or running out of memory.
Diagnose the decision from live objects
Start with the actual numeric priority and class, then inspect pending state, nominated placement, events, and victims:
kubectl get priorityclass
kubectl get pod PENDING_POD -n production \
-o jsonpath='{.spec.priorityClassName}{" priority="}{.spec.priority}{" nominated="}{.status.nominatedNodeName}{"\n"}'
kubectl describe pod PENDING_POD -n production
kubectl get events -n production --sort-by=.metadata.creationTimestamp
kubectl get pdb -n production
kubectl get pods -n production -o wide
The nominated node is evidence of an in-progress scheduling decision, not proof that the Pod is running or that the node has been reserved exclusively. Correlate scheduler events with the victim Pods’ deletion timestamps and termination grace, the controller creating replacements, PDB current healthy and allowed disruptions, and node-level allocated requests. Also inspect quota: if no Pod object was admitted, there is no scheduling decision for priority to influence.
When a high-priority Pod remains Pending, classify the failure before changing values:
- No feasible node even after lower-priority removal indicates hard constraints, capacity on individual nodes, or resource requests that preemption cannot repair.
- Victims selected but the preemptor has not started can be normal while graceful termination or a fresh scheduler evaluation is pending.
- A candidate is feasible but the preemptor is rejected by ResourceQuota or admission; fix the policy or capacity plan, not the class value.
- A victim’s controller cannot replace it because the cluster has no room, its own quota is exhausted, or the original constraints are impossible. Preemption may have moved the outage rather than solved capacity.
- Repeated preemption and rescheduling across the same workloads indicate priority classes or reservations are not aligned with the service recovery hierarchy.
Make the policy testable before production
Document who may create PriorityClasses, which teams may use each class, which classes can preempt, and which workloads are eligible victims. For every high-priority service, test a pending replica under realistic node pressure and confirm that lower-priority workloads recover. Test graceful termination with in-flight requests and durable work, and observe what happens when PDBs would all be violated.
Acceptance criteria should verify that an unknown class is rejected, ordinary workloads keep an appropriate default priority, high-priority resource consumption is bounded, and preemption events are attributable to a named class and workload. Confirm that preempted Pods can be rescheduled after the burst, that cluster autoscaling can satisfy constraints it is expected to solve, and that a no-capacity condition produces an alert rather than an endless cycle.
Priority is best understood as an explicit order of service under contention. It helps the scheduler choose which pending work to favor and can trade lower-priority availability for faster recovery of a higher-priority workload. It cannot make an infeasible Pod feasible, preserve every PDB, guarantee immediate execution, or replace capacity engineering.
Related:
- How the Kubernetes Scheduler Actually Places Workloads
- How to Configure Pod Disruption Budgets in Kubernetes
Sources: