FreeBSD truss: Diagnose Live System Calls Without Guessing
Trace FreeBSD system calls with truss, follow child processes, limit captured strings, and compare live tracing with ktrace and DTrace workflows.
truss(1) traces the system calls made by a process or command. It is often the shortest path from “the application hangs” to an observed sequence such as a missing file lookup, repeated connect(2) failures, or an unexpected permission error. It works by stopping and restarting the traced process through ptrace(2), so tracing changes scheduling and timing. Use it as a focused diagnostic instrument, not as a free, zero-overhead production telemetry stream.
FreeBSD also provides ktrace(1)/kdump(1) and DTrace. Those tools overlap in what they can reveal but have different collection and analysis models. truss is useful when an operator needs a live, readable system-call sequence for one command or a process. ktrace records kernel trace events for later decoding. DTrace supports programmable probes and aggregation. Pick the smallest tool that answers the question, then stop collection as soon as the evidence is sufficient.
Trace a command in a bounded test
Start with a harmless reproduction and direct output to a private file:
umask 077
trace_dir=$(mktemp -d /var/tmp/truss.XXXXXX) || exit 1
truss -o "$trace_dir/truss.out" /bin/echo health-check
sed -n '1,80p' "$trace_dir/truss.out"
The command form starts a new process under tracing. It avoids attaching to an unrelated production process and makes it easier to reproduce the behavior with a controlled environment. The trace may include file paths, user names, network endpoints, and argument strings. Protect it like application logs.
To investigate a failure, use a small input that exercises only the relevant code path. For a configuration lookup, run the program with a test configuration. For a connection problem, use a non-mutating health probe rather than a command that writes business data. Avoid tracing an entire build, package upgrade, or service tree unless the volume and overhead are understood.
Attach to an existing process only with approval
The -p form attaches to a process by PID:
truss -p 12345
Replace the example PID only after verifying it with ps, procstat, or another process inventory. Attach permissions depend on system policy and process ownership. Tracing can pause and resume the process, expose arguments and behavior, and perturb latency. For a production service, obtain the service owner’s approval, define a short capture window, write to a protected file, and have a stop condition. Do not attach to a process that handles sensitive credentials without understanding what the selected options can record.
Use -f when child processes are part of the suspected failure path:
truss -f -o /var/tmp/worker-tree.truss /usr/local/bin/worker --check
-f follows descendants created by fork-like operations and adds process identifiers to distinguish events. It can greatly increase trace volume. If the parent launches many children or a persistent worker pool, narrow the reproduction or use a tool better suited to aggregation.
Reduce output to the signal
The default trace may be verbose. Read the exact truss(1) options installed on the host before relying on filters; option sets can evolve. For example, the documented -a option displays argument strings passed to execve(2), and -s controls the maximum string length displayed. Captured strings may contain secrets. Do not enable them unless the question requires argument visibility and the output destination is appropriately protected.
The essential analysis is usually not to inspect every successful syscall. Search for repeated failures, high-latency calls, unexpected path prefixes, descriptor reuse, and the first error that causes a later cascade. Preserve a few lines before and after each relevant failure so the call context remains meaningful:
grep -nE 'ENOENT|EACCES|ECONNREFUSED|ETIMEDOUT' /var/tmp/truss.out | head -80
An error code must be interpreted in context. An ENOENT during normal library probing may be harmless; the same error for a required configuration file may be fatal. A connect(2) error can indicate routing, policy, listener, or timing failure. Correlate the trace with application logs, socket state, filesystem permissions, and a client test rather than declaring root cause from one syscall.
Understand perturbation and privacy
Because truss stops and restarts the process to observe system calls, it can alter timing and contention. A race may disappear or become more frequent under trace. Latency measurements from a traced process are not production performance baselines. For production impact, use a short bounded window, compare behavior before and after, and prefer DTrace or application instrumentation when the tracing question can be answered without repeatedly stopping the target.
Trace output is sensitive. Paths can disclose tenant names or data layout, arguments can contain passwords, and syscall buffers may include user data depending on tracing flags and the specific system call. Do not upload raw traces to public issue trackers. Restrict file permissions, record collection time and command, redact copies carefully, and rotate secrets if they were captured.
Run truss as the same user as the program when possible. Elevating to root can broaden access and produce more sensitive output. If privilege is required by policy, document the reason, scope, operator, and cleanup. Do not leave a trace process attached after the incident; stop it cleanly and confirm the target continues normally.
Compare with ktrace and DTrace
Use ktrace when a trace file is preferable to a live stream or when you need the kdump decoder’s event-oriented output. It has its own trace-mask and file lifecycle. A stale trace file or an enabled tracing flag can collect more than intended, so clear state and protect output according to its manual. Existing article FreeBSD ktrace and kdump covers that model.
Use DTrace when the question needs aggregation, predicates, stack sampling, kernel probes, or correlation across multiple processes. A DTrace script must be reviewed for probe availability, privileges, buffer behavior, and overhead. “Dynamic tracing” does not automatically mean zero impact or unlimited safe collection.
Do not run all three tools at once as a default troubleshooting pattern. The combined overhead can alter the failure and create unmanageable data. Form a hypothesis, choose one instrument, capture a short interval, and update the hypothesis from actual evidence.
A practical failure workflow
For a process blocked on a file or socket, first record the PID, parent process, command, user, and current state. Then use a safe reproduction under truss if possible. If the incident only happens in the existing process, attach briefly after owner approval. Inspect the first failed open, stat, connect, or poll path; confirm it using independent tools such as sockstat, fstat, procstat, or a service-native diagnostic.
If the trace shows repeated retries, determine whether the application has exponential backoff or is spinning. Do not assume a tight repeated syscall is the cause; it may be a symptom of a downstream timeout. If the trace ends unexpectedly, check whether the program exited, was killed, changed process identity, or spawned an untraced descendant.
Clean up the trace file according to policy and ensure no tracer remains. Record the exact command line without secrets, FreeBSD release, duration, relevant syscall excerpt, hypothesis, independent corroboration, and resulting fix. Keep a sanitized excerpt separate from the raw file.
Acceptance criteria
A useful trace session has a stated question, bounded scope, protected output, known target PID or deterministic command, a clear stop condition, and a corroborating test. The result should identify a syscall sequence and failure context, not merely a large log. A fix is accepted only after the original workload succeeds without the tracer and the relevant service-level test passes.
Related:
- FreeBSD ktrace and kdump: Reading Process Behavior from Kernel Records
- FreeBSD procstat: Process Inspection Beyond ps and fstat
Sources: