Skip to content
LinuxDeep Dive Published Updated 8 min readViews unavailable

Linux EEVDF Scheduling: Lag, Virtual Deadlines, and Latency Tradeoffs

Explain Linux EEVDF fairness through lag and virtual deadlines, then measure scheduler behavior without mistaking nice values for real-time guarantees.

The Linux fair scheduler decides which runnable normal-policy thread should run next when several eligible threads compete for a CPU. EEVDF, short for Earliest Eligible Virtual Deadline First, describes a scheduling approach that combines proportional fairness with virtual deadlines. It is not the same as Linux’s SCHED_DEADLINE policy, and a task’s EEVDF virtual deadline is not an application deadline measured in wall-clock time.

The Linux kernel documentation describes the transition toward EEVDF as beginning in kernel 6.6, with implementation details continuing to evolve. Treat the exact algorithm, exposed tuning controls, and diagnostics as kernel-version-specific. A distribution can backport scheduler changes or ship an older branch with different behavior. Inspect the running kernel and its matching documentation before attributing a latency change to one scheduler revision.

Fairness uses virtual service, not equal wall time

The fair scheduler aims to divide CPU service according to task weights. Nice values influence those relative weights: a task with a more favorable nice setting is entitled to a larger share under contention, but a nice value is not a reservation and does not guarantee a response-time bound. If only one task is runnable on a CPU, it can generally use the available CPU regardless of its relative weight.

EEVDF assigns runnable entities a virtual runtime and derives a lag that represents whether an entity has received more or less than its fair share of service. An entity with nonnegative lag is eligible to run. Among eligible entities, the scheduler selects the one with the earliest virtual deadline. A shorter requested slice can produce an earlier virtual deadline and can improve responsiveness for latency-sensitive work while preserving the long-run fairness model.

“Eligible” is a scheduler accounting condition, not a promise that a task will run immediately. The thread must be runnable, allowed on the CPU, in the relevant scheduling class, and not blocked by a higher-priority task or CPU quota. The task can also be delayed by interrupt handling, virtualization, CPU contention, or its own synchronization. This is why a short interactive task can still see high tail latency on an overloaded host.

Virtual deadlines are not real-time deadlines

EEVDF’s deadline is measured in virtual scheduler time and ranks eligible fair-class work. SCHED_DEADLINE is a separate scheduling policy with runtime, relative deadline, and period parameters and an admission-control model. An application that must satisfy a real-time contract should not assume that nice values or EEVDF time slices provide the necessary guarantee. It must analyze the appropriate real-time policy, kernel configuration, CPU isolation, interrupt load, and failure behavior.

The fair scheduler also has to balance work across CPUs. Per-CPU run queues, affinity masks, topology, and load balancing affect which CPU an entity can use and which competing tasks it sees. CPU affinity may improve cache locality but can also concentrate runnable work onto too few cores. EEVDF’s fairness accounting cannot make an impossible CPU placement meet a latency objective.

Resource controls sit alongside this scheduling logic. A cgroup CPU weight affects relative allocation between groups under contention, while cpu.max imposes a bandwidth limit. A task can be eligible according to the fair scheduler and still be throttled by its cgroup quota. Investigate per-thread scheduler state and the service’s cgroup counters before deciding that the scheduler algorithm is at fault.

Sleeping tasks and lag decay

The scheduler must account for a task that sleeps and later wakes. If it simply reset the task’s lag on every short sleep, a workload could repeatedly block briefly and gain unfair service over continuously runnable tasks. The current EEVDF documentation describes a deferred-dequeue and lag-decay mechanism: a sleeping entity can remain represented on the run queue while its lag decays over virtual runtime, after which its lag is reset.

The details of that mechanism are implementation-sensitive and may change as the kernel evolves. Do not build userspace policy around internal run-queue states or assume that a particular sleep duration buys a fixed scheduling advantage. If wake-up latency matters, measure the target kernel under representative competitors and look at scheduler trace events rather than reverse-engineering one code snapshot into an ABI promise.

Wakeup behavior also depends on where a task wakes, whether it is allowed on the CPU, cache and NUMA topology, and whether a sibling thread is running. A thread can be runnable yet not executing because another eligible entity has an earlier virtual deadline. This is ordinary scheduling competition, not necessarily a lost wakeup or broken timer.

Inspect the actual workload and kernel

Start with the scheduling policy, nice level, CPU affinity, and cgroup placement of the affected thread. A process may have different policies or masks across threads. On Linux, /proc/PID/sched exposes scheduler-related state, while /proc/PID/status includes affinity information. chrt -p reports policy and priority for a task. /proc/schedstat can provide aggregate scheduler statistics when the kernel exposes them.

uname -r
ps -L -o pid,tid,cls,rtprio,ni,psr,stat,comm -p "$PID"
chrt -p "$PID"
sed -n '1,100p' "/proc/$PID/sched"
cat "/proc/$PID/cgroup"

These are read-only observations. A shell command that names a process leader may not show the properties of every worker; use the thread IDs from ps -L and inspect the relevant task path under /proc/PID/task/TID. A container can have a different cgroup namespace view from the host, so collect both views when possible.

For event-level evidence, use scheduler tracepoints or perf sched timehist over a short, bounded workload. Record runnable-to-running delay, migrations, wakeups, context switches, CPU utilization, and cgroup throttling. Trace collection can perturb timing, so compare instrumented and uninstrumented runs. Scheduler traces identify timing and transitions; they do not alone explain why the application needed that CPU time.

Latency controls and version caveats

The EEVDF documentation discusses requesting time slices through scheduler attributes. In the current upstream implementation, sched_attr.sched_runtime is interpreted as a requested slice for a fair-class task and clamped by the scheduler; for SCHED_DEADLINE, the same field is deadline runtime. The UAPI header still documents that field under SCHED_DEADLINE, so the exact attribute behavior and availability are kernel-version details. Do not assume that any Linux release accepts an arbitrary fair-class slice request or that the command-line nice utility controls it. Probe the target API, read the matching headers and manual page, and handle EINVAL, EOPNOTSUPP, and permission failures explicitly.

Before changing scheduler attributes, check whether the latency problem is caused by CPU quotas, affinity, blocking I/O, lock contention, interrupts, or frequency limits. Reducing a time slice may improve one task’s responsiveness while increasing context switching or reducing throughput. Raising priority can starve other work. The correct experiment compares service-level tail latency and throughput under the same load, not a synthetic context-switch counter alone.

For hard real-time requirements, evaluate the real-time scheduler policy and its admission/privilege requirements separately. A low-latency EEVDF task is still part of the fair scheduling class and does not receive a hard deadline guarantee. For normal services, cgroup weights, CPU quotas, and reasonable affinity boundaries are often safer operational tools than setting custom per-thread scheduler attributes.

A reproducible scheduler experiment

Build a test that has one latency-sensitive worker and a fixed number of CPU-bound competitors. Run a baseline with the production policy and cgroup configuration. Then vary only one input: nice value, CPU affinity, cgroup weight, quota, or a supported scheduler attribute. Keep the CPU set, workload, frequency policy, background activity, kernel, and test duration stable. Collect p50/p95/p99 wake-to-run and request latency along with throughput and throttling counters.

Repeat the test after reboot or on another host of the same fleet class if deployment behavior matters. A VM may hide physical scheduling and topology effects; a container shares the host scheduler and inherits its kernel version. Keep raw trace data for representative runs and include the exact commands and feature-probe results in the test record.

When a regression appears after a kernel upgrade, compare scheduler documentation and source for both kernels, but also verify firmware, microcode, CPU frequency, cgroup configuration, and workload mix. Scheduler changes can alter timing without violating fairness goals. The right result is an explained and measured change, not an assumption that “new scheduler” must be faster or slower.

Acceptance criteria

An operational review should state the kernel release and configuration, scheduler policy, nice values, thread affinities, cgroup CPU weights and limits, relevant CPU topology, and instrumentation used. The workload should meet its latency SLO under the expected number of competing runnable tasks without starving background work or breaching throughput requirements. Include an overload case and an affinity-restriction case.

Do not accept a scheduler tuning change based solely on the process’s average CPU usage or a single perf sample. Verify task-level run delay, service-level tail latency, work completed, cgroup throttling, and fairness to other services. Keep rollback settings and ensure the service manager applies the intended controls consistently after restart.

EEVDF provides a way to reason about fair scheduling and latency tradeoffs, not a universal fix for delayed work. Use scheduler accounting to generate a testable hypothesis, then validate it against the actual kernel and workload. If the task is blocked on I/O or a lock, changing its virtual deadline cannot make the missing dependency complete sooner.

Related:

Sources:

Comments