Linux SCHED_DEADLINE: Runtime Budgets, Admission, and Missed Deadlines
Configure and diagnose Linux SCHED_DEADLINE by relating runtime, relative deadline, period, bandwidth admission, affinity, and end-to-end timing limits.
Linux SCHED_DEADLINE is a real-time scheduling policy for threads that have a declared execution budget and recurring timing requirements. It uses earliest-deadline-first scheduling with a Constant Bandwidth Server (CBS) mechanism to account for execution. It is not a setting that makes an arbitrary application finish by a wall-clock deadline: the kernel can schedule eligible CPU execution, but it cannot make a blocked device, overloaded service, unbounded lock holder, or remote dependency respond in time.
Keep it distinct from the fair scheduler’s EEVDF virtual deadlines. EEVDF selects among normal-policy tasks according to fairness and virtual service; SCHED_DEADLINE is an explicit real-time policy with a runtime, deadline, and period. Its parameters create a reservation-like CPU demand that must be admitted by the scheduler, and careless values can reject a policy change or starve other workloads.
Interpret runtime, deadline, and period
A deadline task declares three time values in nanoseconds: runtime Q, relative deadline D, and period T. The intended model is that the task should receive up to Q units of execution during each T interval, with that budget available within D from the beginning of the interval. A normal constrained configuration satisfies 0 < Q <= D <= T. These are CPU execution and scheduling parameters, not a promise that a function call completes in D nanoseconds of elapsed time.
For a task with runtime 2 ms and period 10 ms, the configured utilization is 0.2 of one CPU in the simple single-task model. The scheduler assigns a scheduling deadline consistent with the CBS rules and orders runnable deadline tasks by their scheduling deadlines. If the task consumes its budget before replenishment, it can be throttled until its next eligible period. That behavior is intentional accounting, not necessarily a scheduler hang.
The relative deadline can be shorter than the period when the work must complete early in a recurrence. It can also equal the period for a simpler periodic model. Do not set a very long period merely to make the calculated utilization appear small: a long period changes the replenishment and latency behavior. Select values from measured worst-case execution demand and the application’s actual release pattern, then validate missed deadlines under realistic contention.
Use the scheduler ABI deliberately
The Linux sched_setattr() interface carries a struct sched_attr. For SCHED_DEADLINE, the relevant members are sched_policy, sched_runtime, sched_deadline, and sched_period; the three times are expressed in nanoseconds. The exact structure is size-versioned, so callers should initialize its size member and use the documented ABI rather than assuming every kernel has an identical header layout. A policy change can fail with EINVAL for inconsistent parameters, EPERM when the caller lacks the required authority, or an admission-related error when the requested bandwidth cannot be accepted.
struct sched_attr attr = {
.size = sizeof(attr),
.sched_policy = SCHED_DEADLINE,
.sched_runtime = 2 * 1000 * 1000,
.sched_deadline = 8 * 1000 * 1000,
.sched_period = 10 * 1000 * 1000,
};
/* Apply to the calling thread; check and report the syscall result. */
if (sched_setattr(0, &attr, 0) == -1)
perror("sched_setattr");
This is an ABI sketch; include the Linux scheduler headers available for the target libc and kernel, define the syscall wrapper where the libc does not provide one, and handle structure-version differences. A successful syscall means the policy was installed, not that the workload meets its end-to-end service-level objective. Record the effective thread ID, parameters, affinity, and kernel version alongside test results.
Understand bandwidth admission and CPU domains
The scheduler checks whether a requested deadline bandwidth can be admitted within the relevant scheduling domain. A useful first approximation is the sum of each task’s Q/T, but production feasibility also depends on the kernel’s admission rules, root-domain topology, CPU affinity, bandwidth limits, and any configured deadline bandwidth controls. Do not use a single aggregate utilization calculation as proof that every task is feasible.
Affinity is part of the scheduling design. A task pinned to one CPU competes within that CPU’s eligible scheduling capacity; a task allowed across a set of CPUs is not necessarily interchangeable with several independent reservations. Changing a thread’s affinity, moving it into a cpuset, or changing CPU hotplug state can alter which CPUs are eligible. Check the actual mask and cpuset placement after deployment instead of assuming that a process-level setting reaches every thread.
Admission can protect the system from accepting more deadline work than its configured model permits, but it does not account for every source of end-to-end delay. Interrupt handling, non-preemptible sections, firmware stalls, memory pressure, page faults, I/O, lock contention, and external calls can all consume the application’s response-time budget. Measure the complete path, not only CPU service.
Treat throttling and overruns as evidence
When a task uses its runtime, CBS accounting can defer it until replenishment. If this happens on a critical path, first establish whether the declared runtime reflects actual worst-case CPU demand. A task that repeatedly overruns its budget may have an undersized runtime, unbounded work, a change in input size, or an incorrect period model. Increasing Q without rechecking admission can make the system less schedulable.
Instrument release time, start time, completion time, execution time, and deadline miss count in the application. Include the thread’s scheduling policy and effective parameters in diagnostics. Correlate those observations with scheduler tracing and CPU utilization, while keeping tracing overhead in mind. A high CPU percentage alone does not show whether the task is missing its deadline; a low percentage does not prove it is not blocked on a lock or I/O.
If the workload occasionally needs more execution than its nominal budget, examine whether deadline bandwidth reclaiming is appropriate for the target kernel and workload. Reclaiming is an optional policy mechanism, not extra guaranteed capacity. Test behavior when competing tasks are busy and when idle capacity disappears. The safe baseline is to configure and validate a feasible budget without assuming reclaiming will rescue a miss.
Avoid turning a CPU reservation into a false end-to-end guarantee
Real-time priority can worsen system responsiveness if applied to a task that blocks while holding a lock, busy-loops, or starves helper threads. Analyze lock ownership and priority interactions, isolate unrelated work where required, and define a failure policy for overload. A hard timing objective also needs a defined response to a deadline miss: degrade the feature, discard stale work, restart a bounded operation, or enter a safe state according to the application domain.
Do not migrate a complex multithreaded process to SCHED_DEADLINE based on one benchmark. The scheduler policy is per thread, so setting one thread does not automatically set every worker. Identify which threads need the policy, ensure they do not wait on work scheduled at a lower urgency without a protocol, and check whether helper threads, completion handlers, and device interrupts can meet the same timing path.
Validate with controlled load
Start with a non-production host or a reversible service configuration. Confirm that the kernel supports the policy, inspect all relevant thread parameters, and test the requested admission under the actual affinity and cpuset arrangement. Exercise typical and worst-case input, CPU contention, I/O delay, lock contention, and CPU hotplug or service restart where those events are in scope. Track deadline misses and tail latency, not just average throughput.
Compare results before and after the policy change using the same kernel, hardware, workload, and observability setup. Retain an ordinary-scheduler fallback and a rollback path. If enabling deadline scheduling changes the result, determine whether the gain came from lower dispatch latency, changed CPU allocation, or a test artifact. Avoid treating benchmark success under idle conditions as evidence of production feasibility.
SCHED_DEADLINE is useful when a thread’s recurring CPU demand can be stated and measured. Define the budget honestly, account for the CPUs the thread may run on, observe throttling and misses, and keep the application’s end-to-end deadline separate from the kernel’s CPU scheduling contract.
Related:
- Linux EEVDF Scheduling: Lag, Virtual Deadlines, and Latency Tradeoffs
- PREEMPT_RT Real-Time Patches Merge Into Mainline Linux
Sources: