Skip to content
LinuxDeep Dive Published Updated 7 min readViews unavailable

Linux Timekeeping: Clock Sources, Clock Events, and Scheduler Time

Understand Linux timekeeping by separating the system timeline, interrupt-generating clock events, scheduler timestamps, and delay mechanisms.

Linux timekeeping is not one clock. The kernel combines hardware counters, interrupt sources, scheduler timestamps, timers, and architecture-specific delay mechanisms. A clocksource answers “where are we on a timeline?” A clock-event device generates an interrupt when a timer should fire. sched_clock() supplies fast scheduler-oriented timestamps, and delay timers support short waits where ordinary scheduling is inappropriate. Confusing these roles can produce faulty driver assumptions, misleading benchmarks, or an attempted fix that merely changes which clock a subsystem uses.

The exact devices and paths depend on architecture, kernel configuration, and virtualized hardware. Start with the kernel timekeeping documentation for the target system, then inspect the runtime selection and clocksource driver rather than assuming a timer name from another machine applies.

Clocksource: a continuous counter

A clocksource is the kernel’s basis for tracking elapsed time. It must have documented frequency and characteristics such as width, read cost, and reliability. The kernel converts counter deltas into time units and handles wraparound according to the counter mask and multiplier/shift representation. A higher nominal frequency does not alone mean a better clock: stability, synchronization across CPUs, virtualization behavior, and read overhead matter too.

The kernel can expose available and current clocksource names through sysfs on configured systems. These files are useful diagnostics, not a license to change the active source without understanding why it was selected. A virtual machine’s counter may pause, drift, or be scaled by the hypervisor; a multi-socket platform may have counters that are not perfectly synchronized.

Clocksource quality is a tradeoff among precision, stability, read latency, power, and the platform’s ability to keep counters synchronized. The kernel’s watchdog can compare candidate clocks and mark an unstable source, but a watchdog report needs to be interpreted alongside firmware and virtualization evidence. A clock that is accurate on one host generation may behave differently after a firmware update or migration to another machine class. Preserve dmesg around boot and resume, not just the current source string.

Counter conversion also explains why raw readings should not be compared casually. A hardware counter’s units are not necessarily nanoseconds; the kernel applies conversion parameters and wraparound handling. Sampling two CPUs may add migration effects, while measuring with a userspace system call may include scheduling and cache noise. Prefer kernel timekeeping interfaces for kernel code and documented clock APIs for user space rather than reading a device counter directly.

cat /sys/devices/system/clocksource/clocksource0/available_clocksource
cat /sys/devices/system/clocksource/clocksource0/current_clocksource

The paths may be absent on systems without the corresponding configuration. Do not treat a missing sysfs file as proof that timekeeping is broken. Check kernel configuration and architecture documentation. Capture the current source before comparing systems, and do not force a boot parameter as a first-line performance tweak.

Clock-event devices: interrupts and timer delivery

A clock-event device schedules interrupts or timer events. It may be periodic or operate in one-shot mode, and the selected device can vary by CPU. High-resolution timers rely on appropriate clock-event support and kernel configuration. A clocksource and clock-event device may be physically related on one platform but remain separate abstractions: one reads elapsed time, the other requests a future interrupt.

Tickless operation changes when periodic scheduler ticks are generated; it does not eliminate all interrupts or make timers exact at an arbitrary wall-clock instant. Timer expiration indicates that the kernel may run the timer callback. Scheduling delays, interrupt masking, power states, and load can affect when a userspace thread actually runs. A timeout is normally a deadline or eligibility point, not a real-time guarantee.

Timer resolution and timer accuracy are separate properties. A high-resolution timer API can represent a fine-grained expiry, yet the callback may execute later because the CPU is busy, interrupts are delayed, or the machine is in a power state. Conversely, a periodic source can still satisfy a coarse timeout if the workload allows slack. When debugging a missed deadline, measure the requested expiry, the timer interrupt, callback start, runnable time, and actual task execution as separate timestamps.

For a driver, use the appropriate kernel timer API rather than programming a hardware clock-event device directly unless the driver owns that hardware contract. If a callback can sleep, defer it to a valid process context. Keep hard-interrupt work bounded and understand whether the chosen timer runs in hardirq, softirq, or threaded context for the kernel version.

Scheduler timestamps and delay loops

sched_clock() is used for scheduler timing and tracing. Its requirements and implementation can differ from the system timekeeping clocksource. It may be optimized for fast reads and need not provide the same cross-CPU or suspend semantics that user-visible timekeeping APIs promise. Do not use it as a replacement for a documented clock API simply because it is cheap.

Delay timers support short, busy-wait delays in contexts that cannot sleep. Busy waiting consumes CPU and should be limited to the precise hardware sequencing requirement. For longer waits in process context, use a sleepable timing mechanism so the scheduler can run other work. Replacing a microsecond-level required delay with a sleep can be incorrect; replacing a millisecond busy loop with udelay() can waste CPU and harm latency.

Wall time, elapsed time, and suspend

User-visible calendar time is not interchangeable with elapsed duration. Real-time clocks can be adjusted by NTP, an administrator, or the user. Monotonic clocks are intended for measuring elapsed intervals and do not jump when wall time is corrected, but variants differ in suspend accounting. Kernel and userspace APIs expose these semantics explicitly; choose based on whether the requirement is “at 10:00 UTC,” “after 30 seconds of active runtime,” or “after 30 seconds including suspend.”

Never measure a benchmark by subtracting wall-clock timestamps if the system clock can step during the run. Use an elapsed-time API with the required suspend semantics. For distributed deadlines, account for clock synchronization error and network delay; local monotonic time cannot be compared directly across machines.

Suspend, virtualization, and source changes

When the system suspends, hardware counters can stop or continue, and the kernel must account for elapsed time according to the platform. Some clocksource implementations require reinitialization or correction around suspend. A time jump after resume may therefore be a platform, firmware, hypervisor, or kernel issue rather than an application arithmetic bug.

Virtualization adds another layer: the hypervisor may provide a paravirtualized clock, scale a counter, or emulate timer interrupts. CPU migration between hosts or live migration can expose clock differences. Diagnose with kernel logs, current and available clock sources, hypervisor documentation, guest tools, and repeated measurements across suspend/resume or migration. Do not infer a root cause from a single date sample.

A disciplined timekeeping investigation

Start by describing the symptom with the required semantics: clock goes backward, deadline runs late, CPUs disagree, timestamps drift after suspend, or a short delay is inaccurate. Then collect kernel and platform version, architecture, hypervisor type, selected clocksource, timer configuration, and relevant logs. Compare monotonic elapsed measurements with real-time clock readings to separate wall-clock correction from timer delivery.

Use tracing to compare timer expiration, interrupt delivery, callback execution, and task scheduling. A userspace callback running late does not prove the clocksource itself is wrong; the event may have fired on time while the task waited for CPU. Conversely, a counter source can drift even if wakeups appear punctual over a short test. Repeat measurements under idle and load, on each relevant CPU or NUMA node, and across resume or migration where applicable.

For tests that span minutes or hours, compare against an independent reference only when its synchronization error is understood. NTP correction, leap-second handling, and virtualization clock discipline can alter real-time timestamps without implying that monotonic elapsed time is wrong. Use repeated samples and report distributions rather than one difference. State the exact clock API, suspend policy, kernel version, CPU placement, and hypervisor in the test record so another engineer can reproduce the result.

Avoid applying clocksource= boot parameters or disabling power states based on folklore. Such changes can hide a firmware defect, reduce power efficiency, or break guests that depend on a particular paravirtual clock. Establish a reproducible before/after measurement and validate the documented device contract.

Linux timekeeping becomes understandable when timeline measurement, timer interrupt generation, scheduler timestamps, and busy delay are treated as distinct mechanisms. Diagnose the layer that violates the required time semantics, select the clock API or kernel primitive that matches the requirement, and verify behavior on the actual architecture and platform.

Related:

Sources:

Comments