Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux perf sched timehist: Measuring Wakeup-to-Run Delay

Use perf sched traces to separate runnable queue delay from execution time, inspect scheduler latency, and build bounded, reproducible workload evidence.

When a Linux task feels slow, the delay may be in the application, blocked on I/O, runnable but waiting for CPU, or delayed by a scheduling and wakeup path. perf sched records scheduler tracepoints and provides reports that help distinguish time a task spent waiting to run from time it spent on CPU. That evidence is useful for latency investigations, but it is not an application profiler and does not automatically explain why a task became runnable late or what it did after it ran.

The central quantity in a wakeup-to-run investigation is run-queue delay: the interval between a task becoming runnable and being scheduled on a CPU. It differs from off-CPU time, which can include sleeping, blocked I/O, lock waits, and deliberate waits. It also differs from end-to-end request latency, which includes userspace work, queues, network delays, and dependencies. Report which interval is being measured and avoid labeling every long response as scheduler latency.

Verify perf and kernel trace support

Check that the perf binary matches or is compatible with the running kernel and that scheduler tracepoints are available. Distribution packages may separate the perf utility from the kernel image:

uname -r
perf version
perf list 'sched:*'

The ability to list events does not guarantee permission to record them. Kernel policy, perf event restrictions, container boundaries, and security settings can limit access. Run with the least privilege that permits the capture and follow host policy; do not permanently weaken perf_event_paranoid merely to complete one test. Record the exact perf version, kernel release, privilege context, and event set used.

Record a bounded workload

Use a controlled reproduction rather than leaving a system-wide recording running indefinitely. A command-scoped example records scheduler activity while a test command runs:

sudo perf sched record -- ./reproduce-latency.sh

Replace the script with a deterministic workload whose duration and load are known. If the issue occurs in a service, a command-scoped capture may not include unrelated service processes; use the documented system-wide or PID-scoped options for the installed perf sched version when broader capture is required. Keep the capture short, note start and stop times, and avoid recording sensitive workloads without approval.

After the recording, inspect the time history and latency summaries:

sudo perf sched timehist
sudo perf sched latency

The available reports and columns vary across perf releases. Consult perf sched --help and the local perf-sched(1) manual before using advanced filters or sort keys. Preserve the raw perf.data file and command line before experimenting with report options. A report generated from a different version may format fields differently.

Interpret the time history correctly

perf sched timehist organizes task scheduling events into intervals that help show when tasks were scheduled, how long they ran, and how long they waited after becoming runnable, subject to the event data captured and report options. In the report terminology, distinguish scheduler delay (wake-up to the next schedule-in) from the broader interval between a task scheduling out and scheduling in again; the latter can include time blocked or sleeping before the task is runnable. Runtime is time scheduled on a CPU, not the whole elapsed request time. A long scheduler-delay interval can be evidence of CPU contention, affinity restrictions, priority interactions, cgroup CPU controls, or scheduler placement constraints. It does not prove which of those caused the delay.

Compare task identifiers, CPU assignments, wakeup events, and runtime across repeated samples. If a task wakes on one CPU and runs on another, account for affinity, migration, and topology. If the trace shows long periods when the task is not runnable, investigate blocking syscalls, locks, timers, and I/O instead of attributing that time to scheduler queueing. If the trace begins after the request entered the application, it cannot explain earlier queueing or network delay.

Latency reports aggregate observations. An average can hide a small number of severe outliers, and a maximum can be dominated by one unrelated event. Pair summaries with individual time-history intervals and workload timestamps. Use percentiles only when the tool or your analysis computes them from a sufficiently representative sample; do not infer p99 from a handful of samples.

Correlate with resource controls and application markers

A runnable task can wait because another workload consumes CPU, because its CPU set is narrow, because it is subject to cgroup controls, or because scheduler priority and class affect selection. Record the process’s CPU affinity, scheduling policy and priority, cgroup path, cpu.max, cpu.weight, cpuset restrictions, and competing load. A low task runtime alongside high run-queue delay suggests a different problem from high runtime, but the distinction should be confirmed with the application and trace.

For a cgroup-managed service, inspect its effective hierarchy rather than only the leaf settings. An ancestor quota can constrain the service even when the leaf appears unrestricted. Compare observed delay during both normal load and a controlled contention test. Do not change nice, real-time policy, cpusets, or quotas as an exploratory first response; these can starve other tasks or change system behavior.

Application markers help align scheduler evidence with user-visible requests. Use trace markers or application timestamps only when their semantics and clocks are understood. Clock domains can differ, and writing markers may itself affect timing. Keep one monotonic-time source for application interval measurements where possible, then correlate using documented timestamp conversion rather than matching wall-clock strings by eye.

Avoid misleading conclusions

Scheduler traces show when the scheduler’s observable events occurred, not the full reason a task was delayed. Interrupt activity, virtual-machine steal time, CPU frequency changes, thermal throttling, firmware behavior, and tracing overhead can influence observed execution. A guest’s scheduler view may not include the hypervisor’s scheduling decisions. For virtual machines, inspect host-level scheduling or hypervisor metrics as well.

Tracepoint loss, buffer overruns, a too-short capture, or incomplete process scope can omit the important event. Check perf’s warnings and capture metadata. If the trace reports no delay but the application remains slow, broaden the investigation to blocked time, system calls, filesystem/network dependencies, and application queues rather than declaring the report proof of good latency.

Tracing can perturb a latency-sensitive workload. Compare a baseline without recording to a short capture under equivalent load. Reduce event scope and duration if overhead is material. Do not benchmark production service-level objectives using a trace capture as though it were a passive observer.

An operational investigation pattern

  1. Define the user-visible interval and identify the task or service that owns it.
  2. Record kernel/perf versions, CPU topology, affinity, scheduling policy, cgroup settings, and baseline load.
  3. Capture a bounded, reproducible workload with the narrowest scope that includes the task.
  4. Inspect run-queue delay separately from runtime and blocked intervals.
  5. Correlate outliers with CPU, wakeup source, competing tasks, and application timestamps.
  6. Change one control in staging, repeat the same capture, and check both target and collateral workloads.

Acceptance should be based on a repeated workload: the source of the delay is supported by trace evidence; the proposed change reduces the relevant interval across representative runs; application response latency improves; no other workload regresses beyond its budget; and a capture-off comparison shows the fix is not merely an instrumentation artifact. Keep raw recordings with retention limits because traces can contain process and workload metadata.

perf sched timehist is a way to inspect scheduler timing, not a verdict on application performance. Use it to separate wakeup-to-run delay from execution and blocking, then combine that distinction with cgroup, CPU, and application evidence. A precise interval and reproducible capture are more valuable than a broad trace with a dramatic but ambiguous maximum.

Related:

Sources:

Comments