Kubernetes Scheduler Profiles: Plugin Tuning and Workload Routing
Use kube-scheduler profiles to tune plugin behavior and route selected Pods safely, while avoiding pending workloads, hidden affinity, and queue conflicts.
Kubernetes scheduler profiles let one kube-scheduler process apply different plugin configurations to different Pods. A Pod selects a profile by setting .spec.schedulerName; that profile can change which scheduling plugins run and how they are configured. This is a focused customization point for workloads that need different placement policy without deploying an entirely separate scheduler implementation.
Profiles do not create a separate control plane or bypass the ordinary scheduler stages. They still participate in filtering, scoring, reservation, and binding, and every profile in one scheduler process shares the same pending-Pod queue. Use them when the default scheduler is suitable but a known workload class needs controlled plugin differences. Deploy a separate scheduler only when it must own an independent queue or implement behavior that the in-process plugin framework cannot provide.
Profile selection is an explicit workload contract
A profile has a unique schedulerName. A Pod whose .spec.schedulerName matches that name is evaluated with that profile. When no name is supplied, the API server defaults the Pod to default-scheduler, so a custom configuration must retain a default-scheduler profile if ordinary Pods are to continue scheduling.
apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
profiles:
- schedulerName: default-scheduler
- schedulerName: latency-aware-scheduler
plugins:
preScore:
disabled:
- name: '*'
score:
disabled:
- name: '*'
This example creates a distinct profile but intentionally disables all pre-scoring and scoring plugins for it. It is illustrative, not a general performance recommendation: removing scoring can change placement in ways that harm utilization or workload locality. In production, keep the existing default behavior unless measurement identifies a specific plugin or weight that should change, then test the effect on representative Pods.
The KubeSchedulerConfiguration API is stable as kubescheduler.config.k8s.io/v1. Store the scheduler configuration where the component owner expects it and pass it with kube-scheduler --config. Do not copy an old v1beta3 example from an older runbook; that configuration version was deprecated and later removed. Confirm the component config API and plugin names against the Kubernetes release running in the control plane.
If a Pod names a profile that is not configured or whose scheduler process is not running, it can remain Pending without ever being bound. The Pod has not failed container startup; no scheduler has accepted responsibility for placing it. Include profile names in deployment templates, cluster admission policy where selection must be restricted, and runbooks. A scheduler name is routing metadata, not an authorization mechanism.
Choose a profile versus a separate scheduler
An in-process profile is usually simpler for a variant of built-in placement behavior. It reuses the component binary, its API and cache behavior, and a shared pending queue. Profiles can have different plugin sets or plugin arguments, but their queueSort plugin and configuration must be identical because there is only one queue to order pending Pods.
A separate scheduler process has a separate queue and can be deployed and versioned independently. It also adds another control-plane component to run, upgrade, secure, monitor, and give API permissions. Every scheduler in a cluster needs a unique name. If replicas use leader election, configure a distinct lock identity and the required permissions; do not accidentally contend with the default scheduler’s leader-election object.
The distinction matters when an operator says “we need a custom scheduler.” If the actual requirement is a node-pool constraint, a profile using NodeAffinity may suffice. If the requirement is a custom scoring rule, prefer a scheduler framework plugin that implements the relevant extension point. If the requirement is an independent queue, custom scheduling algorithm, or special scheduler lifecycle, a separate scheduler may be justified. Document why the chosen model matches the operational need.
Understand plugins before changing their stages
The scheduling framework exposes distinct extension points. Filter rejects nodes that cannot run a Pod; Score ranks nodes that remain feasible; later points reserve resources and bind the Pod. Disabling a filter can change correctness by allowing an invalid placement, whereas changing a score usually changes preference among already feasible nodes. Do not treat every plugin as a tunable optimization knob.
Scoring plugins contribute weighted values, so a weight change changes the balance among that plugin and the other enabled scoring plugins. A high weight does not force a node choice if a hard filter rejects it, and it does not guarantee a specific node if other scores or constraints differ. Profile changes should be evaluated against actual candidate-node sets, plugin metrics, and workload-level outcomes such as pending time, fragmentation, and latency.
Plugin configuration is versioned component configuration. The default plugin set may evolve between Kubernetes releases, and a configuration that disables all plugins for an extension point can silently opt out of behavior added later. Prefer narrow changes, render or validate the configuration for the exact binary, and review release notes before upgrading scheduler profiles.
Apply node affinity to a profile carefully
Kubernetes supports an addedAffinity setting for the NodeAffinity plugin on a scheduling profile. This adds a hidden constraint to every Pod using that schedulerName, in addition to any node affinity in the Pod spec. The example below associates gpu-pool-scheduler with nodes labeled for an accelerator pool:
apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
profiles:
- schedulerName: default-scheduler
- schedulerName: gpu-pool-scheduler
pluginConfig:
- name: NodeAffinity
args:
addedAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: workload.example.com/pool
operator: In
values: [gpu]
Only nodes that satisfy this profile-level rule and the Pod’s own required node-affinity terms can be selected. Label the intended nodes consistently, version the labels with the pool, and verify that resource requests, taints, topology, and storage constraints still permit placement. Because profile affinity is not visible in the Pod spec, annotate the scheduler name’s documented pool mapping and expose the reason in operational dashboards or admission feedback.
This is not a guarantee that the Pod can use a GPU; it only constrains node selection. Device allocation, driver readiness, extended resources or Dynamic Resource Allocation, runtime libraries, and application compatibility remain separate requirements. Similarly, profile-level node affinity is not a substitute for access control. If users must not choose an arbitrary scheduler profile, enforce the allowed names through an appropriate admission policy and keep the node-pool labels under controlled ownership.
DaemonSet Pods are a special case: the DaemonSet controller does not support scheduling profiles in this way. It creates Pods using the default scheduler, which honors affinity expressed in the Pod template. Do not assign a profile name to a DaemonSet template and assume its Pods will receive the profile’s hidden addedAffinity.
Route a workload through a profile
Set the scheduler name in the Pod template, not only on an ad hoc test Pod, so every replica and rollout revision follows the same policy:
apiVersion: apps/v1
kind: Deployment
metadata:
name: image-processor
namespace: workloads
spec:
replicas: 3
selector:
matchLabels:
app: image-processor
template:
metadata:
labels:
app: image-processor
spec:
schedulerName: gpu-pool-scheduler
containers:
- name: worker
image: registry.example.invalid/image-processor:1.2.3
resources:
requests:
cpu: "2"
memory: 4Gi
example.com/gpu: "1"
limits:
example.com/gpu: "1"
This workload is only an example: the image and extended GPU resource name are placeholders. A profile does not create or allocate a device. Confirm that the actual device driver advertises the requested resource and that node capacity, tolerations, storage, and application runtime all match.
Roll out by setting the template in one canary Deployment and observing its Pods. Confirm that the profile exists in the active scheduler config, check Pod events and .spec.schedulerName, and verify the bound node satisfies the intended pool. If the Pods remain pending, separate “no scheduler claimed this Pod” from “the selected scheduler found no feasible node.” The first points to profile naming or scheduler availability; the second points to constraints, resources, plugin outcomes, or node state.
Operate and troubleshoot profile changes
When a Pod is Pending, inspect its events and scheduler identity before changing its resource requests:
kubectl -n workloads get pods -l app=image-processor -o wide
kubectl -n workloads describe pod <pod-name>
kubectl -n kube-system get pods -l component=kube-scheduler -o wide
The scheduler’s events use the Pod’s scheduler name as the reporting controller. In a managed Kubernetes service, scheduler configuration may be provider-owned; use the provider’s supported customization options rather than editing a control-plane manifest that is not yours to manage.
For a configuration rollout, preserve the previous known-good scheduler config, validate the component configuration for the running version, and canary the new profile on a small workload. Monitor pending-Pod age and reasons, scheduler queue and scheduling latency, plugin errors, node distribution, and workload-level performance. Review both the default profile and every custom profile after upgrades because plugin defaults can change.
A profile change is safe only if each profile has a unique name, default-scheduler still exists for ordinary Pods, every custom PodTemplate refers to a live profile, the shared queueSort configuration is compatible, and all added affinity terms match intended nodes. Test deletion and rollback as well as successful placement. When removing a profile, first migrate or stop every workload that still names it; otherwise new Pods can remain Pending.
Production acceptance checklist
- Use an in-process profile for a targeted plugin variation; use a separate scheduler only when separate lifecycle or queue behavior is needed.
- Configure
KubeSchedulerConfigurationwith the API version supported by the exact scheduler binary. - Keep the
default-schedulerprofile and give every additional profile a unique, documented name. - Confirm every Pod template that selects a profile has a running scheduler responsible for that name.
- Keep one compatible
queueSortplugin and configuration across profiles in the same scheduler process. - Change plugin stages and score weights narrowly, understanding that filters affect feasibility while scores affect ranking.
- Treat
addedAffinityas an invisible additional constraint, label nodes consistently, and test interaction with Pod affinity and resource availability. - Do not rely on profile selection as authorization or assume DaemonSet Pods use non-default profiles.
- Canary configuration changes and upgrades; measure Pending age, scheduling latency, chosen nodes, and workload outcomes.
Scheduler profiles provide a manageable way to express deliberate placement differences inside the existing kube-scheduler framework. They are production-ready when the name-to-policy contract is explicit, plugin changes are narrow and versioned, hidden constraints are observable, and a missing profile cannot strand workloads unnoticed.
Related:
- Kubernetes Pod Scheduling Gates: Coordinate Readiness Before Placement
- Kubernetes Topology Spread Constraints: Design for Failure Domains
Sources: