Skip to content
SRE & DevOpsDeep Dive Published Updated 9 min readViews unavailable

Kubernetes Descheduler Operations: Safe Evictions, Node Fit, and Disruption Budgets

Operate Kubernetes Descheduler with bounded eviction policies, node-fit checks, PDB awareness, dry runs, and version-matched configuration.

The Kubernetes scheduler chooses a node when a Pod is first placed. Cluster state changes afterward: nodes are added, labels and taints change, replicas converge unevenly after failures, and resource requests become easier to fit elsewhere. The Kubernetes Descheduler can find selected running Pods that no longer match an operator’s placement policy and evict them so their workload controllers can create replacements.

The boundary is important: Descheduler does not place a replacement Pod. It requests an eviction; the Pod’s owner, such as a Deployment or StatefulSet, decides whether to create another Pod, and the normal Kubernetes scheduler decides where that replacement can run. An eviction can therefore lead to a Pending Pod, a volume-attach delay, or no replacement at all when the object has no suitable controller. Descheduler is a policy-driven disruption mechanism, not a cluster optimizer with a guaranteed move operation. Its upstream documentation describes this division between eviction and rescheduling.

Match the policy to a specific reason to evict

Descheduler has two broad groups of strategy plugins. Deschedule strategies examine candidate Pods individually for a condition such as violating node affinity, tolerating a taint that no longer applies, or exceeding a restart threshold. Balance strategies evaluate groups of Pods against a cluster-wide distribution goal, such as moving replicas off over-utilized nodes when under-utilized nodes exist. A policy may enable several plugins, so an apparently small configuration change can increase the number and type of Pods considered for eviction.

Begin with a concrete desired outcome. If a node was labeled incorrectly and existing Pods violate required affinity, a placement rule strategy may be appropriate. If a workload should be spread according to its existing topology constraint, use the relevant balance policy and verify that the constraint is actually present in the Pod spec. If nodes should be consolidated, ensure there is a separate, tested capacity and disruption plan: a descheduler eviction alone does not cordon or remove a node, and it does not guarantee that a cloud autoscaler will scale one down.

Do not enable every strategy because its name sounds useful. Some policies can repeatedly evict and reschedule a Pod without making progress if the placement preference cannot be met. A soft preference does not reserve an eligible destination. Use a narrowly scoped namespace or label selector when the strategy supports one, and confirm the plugin’s exact filtering semantics for the version in use.

Understand node fit and its limits

The Default Evictor can use nodeFit to test whether a Pod could fit on another node before evicting it. The upstream guide documents checks such as Pod selectors, tolerations and node taints, node affinity, resource requests against available resources, unschedulable nodes, and Pod anti-affinity. This filter reduces avoidable evictions, but it is still a best-effort check against the state and scheduling inputs it understands. It is not a reservation: capacity can change between the check and the replacement scheduling attempt, and a custom scheduler plugin, storage constraint, admission policy, or external controller may add requirements that the filter does not model.

For utilization balancing, the default calculation is based on Pod resource requests, compared with node allocatable capacity; that is not the same as a real-time CPU or memory usage measurement. A node with large requests but low observed usage can still appear utilized to a request-based policy. If actual utilization is the intended signal, configure a supported metrics provider for the exact Descheduler release and verify query semantics, freshness, and failure behavior. Do not interpret kubectl top as proof that a request-based policy is measuring the same thing.

An example conservative policy for a staged trial enables only LowNodeUtilization, limits the number of evictions, avoids Pods with PVCs, and asks the evictor to check that another node could fit a candidate. It deliberately uses illustrative thresholds; there are no universal percentages that are correct for every cluster.

apiVersion: descheduler/v1alpha2
kind: DeschedulerPolicy
maxNoOfPodsToEvictPerNode: 1
maxNoOfPodsToEvictPerNamespace: 1
maxNoOfPodsToEvictTotal: 2
profiles:
  - name: staged-utilization-balance
    pluginConfig:
      - name: DefaultEvictor
        args:
          nodeFit: true
          minReplicas: 2
          podProtections:
            extraEnabled:
              - PodsWithPVC
      - name: LowNodeUtilization
        args:
          thresholds:
            cpu: 20
            memory: 20
            pods: 20
          targetThresholds:
            cpu: 50
            memory: 50
            pods: 50
    plugins:
      balance:
        enabled:
          - LowNodeUtilization

This policy format and supported fields are versioned with the Descheduler application; treat the example as a starting point, not a universal drop-in manifest. Check the policy reference and release branch for the exact image you plan to run. In the LowNodeUtilization example, both threshold maps describe the same resource dimensions, and the plugin may choose candidate nodes based on requests rather than observed consumption. maxNoOfPodsToEvictTotal bounds one rescheduling cycle; per-node and per-namespace bounds further limit how many eligible Pods one run can evict. The PVC protection is an extra guard here, even though persistent-volume behavior should still be tested with the actual CSI driver and workload.

PodDisruptionBudgets are a guardrail, not a complete availability plan

Descheduler uses the Kubernetes eviction subresource, so a Pod covered by a PodDisruptionBudget is not evicted when doing so would violate that budget. A blocked eviction may therefore be correct behavior, not a Descheduler failure. Check the budget’s selected Pods, healthy replica count, disruptionsAllowed, and the workload’s actual availability before changing the policy to make more Pods evictable.

A PDB does not create spare capacity, ensure a replacement can be scheduled, or make application state safe to move. A PDB with an overly broad selector can guard the wrong Pods; a budget with no allowed disruptions can stall maintenance and rebalancing indefinitely. For a stateful workload, confirm that its controller can recreate a replacement, that storage can attach in an eligible topology, and that the application tolerates the interruption. Do not use Descheduler as a way to bypass a deliberate availability budget.

Upstream defaults also protect several categories of Pods from eviction, including system-critical Pods, unmanaged bare Pods, DaemonSet Pods, and Pods with local storage. PVC-backed Pod handling depends on the selected protection settings. Enabling an override can turn a workload that was previously protected into an eviction candidate. Inspect the current release’s protection table before setting options that permit local-storage, DaemonSet, or system-critical evictions. Avoid those overrides during an initial rollout.

Deploy one controlled execution model

Descheduler can run as a Job, CronJob, or Deployment. A Job is useful for a bounded experiment; a CronJob gives operators an explicit, reviewable cadence; a Deployment can run repeated descheduling cycles. Select an execution model that fits the time needed for a replacement to be admitted, scheduled, started, and healthy. Running cycles too frequently can create a burst of simultaneous disruptions while earlier replacements are still Pending or not Ready.

Start with the upstream recommended deployment and RBAC for the exact release rather than copying permissions from an unrelated version. A production policy should also bound total, per-node, and per-namespace evictions, and define whether PVC-backed workloads are eligible. Before enabling more than one replica of a Deployment, configure Descheduler leader election so replicas do not independently process the same cluster at the same time. The upstream high-availability guidance calls out leader election for multi-replica operation.

Run the same image and policy in dry-run mode first. The command-line reference defines --dry-run as executing Descheduler without performing evictions; --policy-config-file selects the policy file. For a container image with the policy mounted at /etc/descheduler/policy.yaml, the invocation is:

/bin/descheduler \
  --policy-config-file=/etc/descheduler/policy.yaml \
  --dry-run \
  --v=4

Review each proposed Pod and the strategy that selected it, confirm the predicted destination set is operationally acceptable, and then test a very small bounded real run in a non-production namespace. Dry run is valuable, but it is not a complete simulation of future scheduler, storage, or admission outcomes. Logs at verbosity 4 or above can explain why candidates were filtered; tune verbosity and retention so the information is available to responders without flooding routine logs.

Observe the whole replacement path

Measure more than the number of eviction requests. Track Descheduler cycle duration, eviction attempts and failures, candidates skipped by policy, and the age of any replacement Pod that stays Pending or unready. Correlate these signals with PDB status, scheduler events, node capacity, node affinity and taints, volume attachment, rollout status, and application-level availability. A low count of evictions may mean the policy is healthy and finds no candidates, or it may mean every candidate is blocked; the reason matters.

Use the version-specific metric names and labels. The upstream project marks some older metric names deprecated and documents replacements; dashboards should be reviewed when the Descheduler version changes. The metrics endpoint is served through HTTPS with authentication and authorization defaults in the CLI configuration. Scrape it through the cluster’s intended monitoring path rather than disabling endpoint security simply to make a dashboard work.

During an incident, pause the Job schedule or scale the single active Deployment to stop further policy runs while investigating. Do not delete the replacement Pods that are already being created just to make eviction counters settle. Identify the workload owner and current Pod generation, compare current requests with the policy thresholds, inspect the PDB and scheduler events, and validate that a proposed destination is feasible. Roll back the policy change and let the workload controller restore a known-good desired state if the new strategy causes repeated churn.

Roll out policy changes in stages

  1. Baseline: Record node requests and allocatable values, Pod distribution, PDBs, storage classes, and workload replica counts. Identify Pods that must never be moved during the trial.
  2. Review the plugin: Confirm its selection rules, eviction protections, node-fit behavior, and version-specific configuration. Keep the first policy to one strategy and one bounded scope.
  3. Dry run: Execute the exact policy on a representative cluster and inspect candidate Pods, protected Pods, and skipped reasons.
  4. Canary: Permit only a small number of real evictions in a non-production pool or namespace. Verify replacement creation, placement, storage, readiness, and application health end to end.
  5. Expand gradually: Increase the scope or eviction limits only after a complete cycle is within the recovery objective. Keep a fast way to stop additional runs and revert the policy.

Before adoption, verify that the tested Descheduler release corresponds to a supported Kubernetes version and use documentation from that release branch. The project explicitly warns that its master documentation is in development and may not match a published release. For each upgrade, review plugin names, policy fields, default protections, metrics, and compatibility instead of assuming a prior policy is still valid.

Descheduler is useful when there is a clear reason to reconsider existing placements and the workload controllers can safely replace evicted Pods. It is not a substitute for correct initial scheduling constraints, capacity planning, topology design, PDBs, or workload health checks. Treat every eviction as the first step in a multi-controller path, constrain how many can occur, and declare success only after the replacement is running and serving safely.

Related:

Sources:

Comments