Linux ftrace Function-Graph Tracing: A Bounded Kernel Call-Path Workflow
Use ftrace function-graph tracing to inspect selected kernel call paths, bound trace volume, interpret timing, and restore global tracing state safely.
ftrace is the Linux kernel’s tracing framework for examining functions, events, and execution behavior. Its function-graph tracer records function entry and return relationships for instrumented kernel functions, which makes nested call paths visible along with elapsed durations. That view can explain why a syscall or driver operation spends time in a particular subsystem. It is also easy to overwhelm the trace buffer, perturb a workload, or alter tracing state that another operator expects. A useful workflow starts narrow, captures a bounded interval, and restores the previous configuration.
Function-graph output is not a complete performance profile. It traces only what the selected tracer and available instrumentation can observe, and tracing itself has overhead. Reported duration is elapsed time in the traced call path, not automatically CPU-exclusive time or end-to-end user-visible latency. Scheduling, interrupts, preemption, clock configuration, and nested calls affect interpretation. Use it to answer a concrete call-path question, then corroborate with a measurement suited to the full workload.
Locate tracefs and check capabilities
Modern systems typically expose tracefs at /sys/kernel/tracing; older layouts may use /sys/kernel/debug/tracing. Confirm the mount and permissions rather than assuming a path:
findmnt -t tracefs
sudo mount -t tracefs tracefs /sys/kernel/tracing
test -r /sys/kernel/tracing/available_tracers && cat /sys/kernel/tracing/available_tracers
Only run the mount command if tracefs is not already mounted at the expected location. Mounting and writing tracer controls require administrative privileges on most systems. A kernel can omit a tracer or function instrumentation support, so check available_tracers, available_filter_functions, and the running kernel’s configuration before planning a capture.
Before changing anything, inspect current tracing state. Tracefs controls can be global to the tracing instance, and another tool may be using them:
TRACE=/sys/kernel/tracing
cat "$TRACE/current_tracer"
cat "$TRACE/tracing_on"
cat "$TRACE/set_ftrace_filter"
cat "$TRACE/set_graph_function"
If a profiler, observability agent, or another operator owns tracing, do not overwrite its settings. Prefer a dedicated trace instance where the kernel and tools support the needed tracer and controls. Record the original tracer, filters, and tracing state so you can restore them even if the experiment fails.
Filter before enabling the graph tracer
Function-graph traces can generate a large amount of output. Select the smallest entry-point set that can answer the question. Check that candidate functions are available:
grep -E '^(do_sys_openat2|vfs_open|path_openat)$' "$TRACE/available_filter_functions"
Names differ across kernel versions and configurations. Discover a real symbol from the running kernel rather than copying a function name from another release. A symbol absent from the filter list cannot be traced by this workflow. Some functions may be excluded from tracing or have compiler/instrumentation constraints.
Use the set_graph_function filter to trace a chosen function and its nested calls. The following is illustrative; verify each function is present and relevant on the target kernel:
sudo sh -c 'echo nop > /sys/kernel/tracing/current_tracer'
sudo sh -c 'echo > /sys/kernel/tracing/trace'
sudo sh -c 'echo vfs_open > /sys/kernel/tracing/set_graph_function'
sudo sh -c 'echo function_graph > /sys/kernel/tracing/current_tracer'
Avoid enabling function_graph across all functions on a busy production host. A narrow graph root limits volume, but its descendants can still be numerous or frequent. If the goal is a particular task or CPU, use supported PID/CPU filters deliberately and read the tracer documentation for their exact interactions. Filters reduce data volume; they do not make tracing free.
Capture a short, controlled interval
Start with tracing disabled while configuring. Enable tracing only for the bounded operation, and stop promptly:
sudo sh -c 'echo 1 > /sys/kernel/tracing/tracing_on'
# Run one reproducible operation in another terminal.
sudo sh -c 'echo 0 > /sys/kernel/tracing/tracing_on'
sudo cat /sys/kernel/tracing/trace > /tmp/ftrace-function-graph.txt
Use a test workload that represents the suspected path, not a broad stress test unless the question requires it. Record the exact operation, workload state, kernel release, CPU topology, and relevant driver/device. Start with a few iterations and inspect buffer loss or overruns. If the trace truncates, narrow the filter or reduce capture duration before enlarging buffers; a larger buffer can hide an overly broad design and consume more kernel memory.
For continuous consumption, trace_pipe provides a streaming view, but reading it consumes trace records. It is not a stable archive by itself. For a repeatable capture, save the trace and metadata, including control values and timestamps. Use a parser or trace visualization tool only after preserving the raw output and documenting its assumptions.
Read the output as a call tree
Function-graph output uses indentation to show nested calls, and timing annotations indicate durations according to the tracer’s output format and configured options. A long parent duration contains time spent in descendants; adding parent and child durations together double-counts nested time. Compare sibling paths and repeated samples rather than summing every printed duration into a supposed total.
An unexpectedly long function can mean the function itself executed slowly, that it was blocked or preempted while active, or that the selected measurement includes nested work. Inspect the relevant trace options and kernel documentation before making CPU-time claims. Correlate with scheduler tracing, block events, IRQ activity, or application timestamps when you need to distinguish running time from waiting. A trace focused on function calls may not tell you why a wait occurred.
Trace records can include process, CPU, and timestamp context. Use those fields to determine whether different events interleaved across CPUs or tasks. Do not infer causation merely because one event follows another in a printed file; concurrent CPUs and asynchronous work require explicit correlation. A function graph is a temporal observation, not a proof of the complete causal chain.
Control overhead and protect production systems
Tracing can perturb cache behavior and scheduling, especially when many short functions are recorded at high frequency. Measure the application both with tracing off and during a short capture. Do not use instrumented timings as if they were uninstrumented benchmark results. If the performance regression disappears when tracing is enabled, that itself is evidence that the measurement overhead may be significant.
Use the least privilege practical for reading trace output, but assume writes to tracefs are powerful host-level diagnostic actions. Avoid collecting sensitive data from unrelated processes. Store trace files with appropriate permissions and retention, especially if event payloads or function arguments could expose operational details.
When using a trace instance, validate which controls are instance-local and which remain global for the particular tracer on the running kernel. Isolated instances reduce interference but do not eliminate hardware or system-wide measurement effects. Check kernel documentation for instance semantics instead of assuming every file is namespaced.
Restore the previous state and define acceptance
When finished, disable tracing before resetting filters and tracer selection. This example demonstrates a clean baseline only if you recorded that baseline first; it is not safe to blindly replace settings belonging to another tool:
sudo sh -c 'echo 0 > /sys/kernel/tracing/tracing_on'
sudo sh -c 'echo nop > /sys/kernel/tracing/current_tracer'
sudo sh -c 'echo > /sys/kernel/tracing/set_graph_function'
sudo sh -c 'echo > /sys/kernel/tracing/set_ftrace_filter'
Restore the actual prior values when there was an existing owner. Verify current_tracer, filters, and tracing_on afterward. If a capture helper or experiment is interrupted, check for a still-enabled tracer before leaving the host.
An acceptable investigation has a written question, a verified symbol/filter, a short and bounded capture, raw output plus environment metadata, a stated interpretation of timing, and a confirmed cleanup. It should also explain what the trace cannot establish and what independent evidence corroborates the conclusion. If the only result is a huge trace file with no hypothesis or cleanup record, the tracing session was not operationally complete.
ftrace function-graph tracing is most powerful when used as a microscope, not a surveillance net. Narrow the call path, capture only what you need, avoid double-counting nested durations, and restore system state. That discipline yields evidence that can guide a fix without turning diagnostics into another production incident.
Related:
- Tracing System Calls and Kernel Events with bpftrace
- Linux Dynamic Debug: Enable pr_debug Call Sites Without Rebuilding
Sources: