Skip to content
SRE & DevOpsDeep Dive Published Updated 10 min readViews unavailable

Kubernetes Dynamic Resource Allocation: ResourceClaims, Drivers, and Scheduling

Deploy Kubernetes DRA with compatible device drivers, DeviceClasses, ResourceClaims, scheduler-aware allocation, and production troubleshooting checks.

Kubernetes Dynamic Resource Allocation (DRA) gives workloads a structured way to request devices such as GPUs, network adapters, FPGAs, and other specialized resources. Instead of only asking for an integer count of a generic extended resource, a workload can reference a claim whose requirements select device attributes and whose allocation is coordinated with Pod scheduling.

DRA is a framework, not a device driver. Installing Kubernetes v1.35 or newer does not make a GPU, NIC, or accelerator claimable by itself. The device owner must provide a compatible DRA driver, the cluster operator must install and configure it, and workload authors must use the DeviceClass and claim model that driver publishes. Treat those layers as separate acceptance gates.

What DRA adds to device allocation

Traditional device plugins commonly advertise extended resources such as vendor.example/gpu; a container requests a quantity, and the scheduler accounts for that quantity on a node. DRA can express richer device selection and configuration through ResourceClaims, DeviceClasses, and driver-reported attributes. It can also support shared access when the device and driver allow that model. DRA does not make every device shareable, and it does not remove the need to understand the vendor’s runtime, topology, or licensing behavior.

The core DRA feature is stable in Kubernetes v1.35 and uses the resource.k8s.io/v1 API. Current upstream documentation marks prioritized request lists stable in v1.36 and DRA-backed extended-resource allocation stable in v1.37. Workload-level ResourceClaims for PodGroups remain a separate beta feature in v1.37 and are disabled by default. Do not rely on a newer capability just because the core DRA API is stable; verify the exact cluster release, enabled feature gates, and driver support for each optional behavior. Kubernetes v1.37 also stabilizes DRA device taints and tolerations; this guide’s examples do not depend on that optional capability.

Understand the API objects and ownership

Four objects form the basic allocation model:

  • DeviceClass describes a category of devices and the selectors that define it. It is cluster-scoped and usually created by a driver owner or cluster administrator.
  • ResourceSlice reports devices, their attributes or capacity, the pool they belong to, and which nodes can access them. A DRA driver publishes and reconciles these objects; operators should not edit a driver’s ResourceSlices by hand.
  • ResourceClaim records a workload’s request and, after allocation, the selected resource. A workload operator can create a claim directly, or Kubernetes can generate one from a template.
  • ResourceClaimTemplate describes a per-Pod claim. The controller creates a distinct ResourceClaim for the Pod, and that generated claim follows the Pod’s lifetime.

Use a directly managed ResourceClaim when several Pods should refer to the same named claim or when the claim needs a lifecycle independent from one Pod. Actual sharing still depends on device and driver capabilities and on whether the application can safely share that device. Use a ResourceClaimTemplate when each Pod needs its own similarly configured allocation, such as parallel workers that must not share one accelerator instance.

A Pod that references an explicit ResourceClaim needs that claim to exist in the same namespace. A missing claim can leave the Pod pending rather than producing a regular container startup error. An automatically generated claim from a template is tied to its owning Pod; do not treat its name as a durable workload interface.

Allocation lifecycle from driver to container

DRA involves three operational roles. The device owner supplies the compatible driver and its resource model. The cluster administrator attaches devices to nodes, installs the driver, and makes appropriate DeviceClasses available. The workload operator creates a ResourceClaim or template and references it from the Pod. These roles need a shared compatibility matrix because the API object alone cannot prove that firmware, driver, runtime, and application versions work together.

At scheduling time, Kubernetes evaluates the request against driver-managed ResourceSlices. The scheduler filters out devices that do not match the claim or are not accessible from eligible nodes, allocates a matching resource to the claim, and schedules the Pod to a node that can use it. The kubelet then coordinates with the node-side driver to prepare the device for the container. Compatible drivers commonly expose device access through the Container Device Interface (CDI), but the exact runtime and device interface remain driver-specific.

With the default behavior documented upstream, pools and ResourceSlices are evaluated using a first-fit ordering based on their names; pools without binding conditions are evaluated before pools that have them. Drivers using Kubernetes’ DRA helper library may manage slice naming to express their intended ordering. Treat this as an allocation-priority detail, not a substitute for expressing the workload’s actual requirements. Put compatibility constraints in a documented DeviceClass or claim selector instead of relying on names to steer production placement.

Request a device with a claim template

This illustrative configuration requests one device from a pre-existing example-device-class whose driver publishes the shown GPU attribute and capacity. Replace the example class, attribute domain, and application image with values documented by the installed driver. The exact attribute schema is not universal.

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: training-gpu
  namespace: gpu-workloads
spec:
  spec:
    devices:
      requests:
        - name: accelerator
          exactly:
            deviceClassName: example-device-class
            selectors:
              - cel:
                  expression: |-
                    device.attributes["driver.example.com"].type == "gpu" &&
                    device.capacity["driver.example.com"].memory == quantity("64Gi")
---
apiVersion: v1
kind: Pod
metadata:
  name: training-worker
  namespace: gpu-workloads
spec:
  restartPolicy: Never
  resourceClaims:
    - name: accelerator
      resourceClaimTemplateName: training-gpu
  containers:
    - name: trainer
      image: example.invalid/trainer:1.0.0
      command: ["./run-training"]
      resources:
        claims:
          - name: accelerator

The spec.resourceClaims entry connects a Pod-local name to a ResourceClaim or template. The container’s resources.claims entry identifies which claim the container uses. A claim declared on the Pod but not assigned to a container is not an application-visible device request. When several containers need the same allocation, each must reference the intended claim explicitly and the driver must support that use.

This is a schema example, not a ready-to-run GPU manifest: the class and attributes must exist in the cluster, the device capacity must match the claim, and the placeholder image does not exist. Before applying a workload, review the driver’s supported selectors, runtime requirements, image compatibility, and resource-sharing policy.

Plan rollout and scheduling behavior

Start by verifying the feature and driver rather than debugging the application:

kubectl version
kubectl get deviceclasses
kubectl get resourceslices
kubectl get resourceclaims -n gpu-workloads
kubectl describe pod training-worker -n gpu-workloads
CLAIM_NAME=training-gpu-abcde # Replace with the generated ResourceClaim name.
kubectl describe resourceclaim "$CLAIM_NAME" -n gpu-workloads

Confirm that the cluster’s API server, scheduler, controller manager, kubelet, and device driver versions satisfy the vendor’s matrix. Verify that ResourceSlices report the expected device pool and eligible nodes, that the DeviceClass exists, and that the claim’s CEL selectors match attributes actually published by the driver. ResourceSlices can change as hardware health and capacity change; compare the live object with driver logs rather than assuming a manifest is the current inventory.

Do not set spec.nodeName to pin a DRA Pod directly to a node. That bypasses the scheduler’s resource allocation path and can leave the Pod unable to use the claim while occupying ordinary node resources. If one node is required for a test, use a node selector or affinity constraint so the scheduler still evaluates the device request.

Kubernetes does not currently preempt an existing Pod to make room for another Pod that needs the same DRA resource. A high-priority workload can therefore remain pending until a conflicting Pod terminates or the device becomes available. Plan accelerator capacity, parallelism, queueing, and job deadlines with that limitation in mind. Pod priority by itself is not a device-sharing or preemption mechanism.

Troubleshoot by allocation stage

When a Pod does not start, trace the object chain rather than changing the container image first:

  1. Pod accepted but pending: inspect Pod events, node selectors, taints, claim references, and whether the referenced claim exists in the Pod namespace.
  2. Claim exists but is unallocated: compare the claim’s DeviceClass and selectors with current ResourceSlices, accessible nodes, capacity, and device health reported by the driver.
  3. Claim allocated but Pod not bound: inspect scheduler events and node eligibility. Check that the selected device is accessible from a node meeting all ordinary CPU, memory, topology, affinity, and taint constraints.
  4. Pod scheduled but container cannot see the device: inspect kubelet and node-driver logs, driver preparation calls, CDI specification discovery, runtime configuration, device permissions, and the vendor’s container interface.
  5. Pod starts but workload fails: verify the application image, user-space driver libraries, device architecture, firmware, licensing, and the actual device paths or metadata exposed to the container.

The Kubernetes API provides allocation status, but vendors define many device attributes and preparation details. kubectl describe output and events help find the failed stage; node-side driver logs explain whether preparation, mount, CDI injection, or hardware initialization failed. Preserve those logs during a canary because a successful allocation in the API does not prove that the application can execute useful work on the device.

Roll out DRA without obscuring the failure boundary

Use a small canary workload with a known device, a deterministic image, and a representative application check. Record the expected DeviceClass, claim selectors, node set, driver version, runtime, and observed allocation before increasing concurrency. Test Pod replacement and deletion so template-generated claims are cleaned up as expected, then inspect for claims left behind by failed controller or driver cleanup.

For parallel Jobs, calculate demand as a function of active Pods rather than Job count alone. If each worker gets an independent template-based claim, increasing parallelism can request more hardware at once. If workers intentionally share one explicit claim, verify that the driver and application support the sharing mode and that concurrent writes or device contexts remain correct. Test scale-down and retries as well as successful startup; device preparation and release are part of the workload’s lifecycle.

Keep the core API separate from optional capabilities in cluster documentation. Prioritized alternatives, device metadata, extended-resource mapping, device sharing, and PodGroup-level claims have their own version or driver requirements. A migration that works with one accelerator vendor does not prove that another DRA driver interprets the same attributes or preparation behavior identically.

Production acceptance checklist

  • The Kubernetes release supports the resource.k8s.io/v1 API and the required DRA feature set.
  • The chosen device driver is DRA-compatible and versioned with the node runtime, firmware, and application image.
  • DeviceClasses and driver-owned ResourceSlices expose the expected attributes, capacity, and node eligibility.
  • The workload uses a managed ResourceClaim or a per-Pod ResourceClaimTemplate for the correct sharing and ownership model.
  • Pod and container claim references match, and the claim exists in the workload’s namespace when managed directly.
  • Scheduler constraints allow a node with the requested device to run the Pod; the workload does not bypass scheduling with nodeName.
  • Capacity, Job parallelism, lack of DRA preemption, and device cleanup are included in rollout and recovery plans.
  • API allocation success and successful application-level device work are verified as separate acceptance signals.

DRA makes device requirements more expressive and gives the scheduler a role in allocating them, but a production device path still spans a driver, cluster configuration, API objects, scheduler decision, kubelet preparation, runtime injection, and application libraries. Validate every handoff with a canary and preserve the distinction between “claim allocated” and “workload can use the device.”

Related:

Sources:

Comments