Profile Haiku Applications with the Native Sampling Profiler
Use Haiku's profile tool to collect repeatable program-counter samples, interpret uncertainty, export Callgrind data, and verify optimizations.
Haiku ships a native profile tool that samples program counters periodically and reports functions where a thread was observed. It can profile a command it starts, recursively include child threads/teams by default, or profile running teams. It also supports output suitable for Callgrind-compatible viewers. This is a statistical sampling tool, not a trace of every function call and not a replacement for reproducing a workload with measurements.
The most important unit in the tool’s current source is the sample interval: -i accepts microseconds, and the default is 1,000 microseconds (one millisecond). A shorter interval can produce more samples on a fast machine but may make results worse on a slow one. Do not copy examples that label the value milliseconds without checking the target revision; -i 300 means 300 microseconds in current source, not 300 milliseconds.
Start with a reproducible workload
Before profiling, define one scenario that looks slow and a measurable completion condition. Record Haiku revision, architecture, CPU, application version, data set, window dimensions, and relevant settings. Close unrelated workloads where practical and perform one warm-up pass so first-run package loading or cache effects are not confused with the steady-state bottleneck.
The simplest mode profiles a command started by the tool:
profile -i 1000 /path/to/application argument1 argument2
Replace the path and arguments with the exact target command. The tool reports its usage with profile --help. In command mode, let the workload reach the measured state and exit normally so results for its threads are emitted. If the program opens a window, reproduce the scenario, close it normally, and retain the complete Terminal output.
For an interactive app that cannot be launched from Terminal, check the current profile --help options for profiling all teams or the relevant running team. System-wide profiling can include unrelated work and makes attribution noisier, so use it only when needed. Do not add -k automatically; it includes kernel frames and changes the question from “where is my application spending samples?” to a broader user/kernel analysis.
Interpret sample counts as evidence, not exact time
Sampling observes execution at intervals. A function hit many times is more likely to consume CPU than a function with few samples, but the count is not an exact call count and does not directly state wall-clock latency. Blocking on network or disk may make a UI feel slow while a CPU profiler shows little activity because the thread is waiting rather than executing.
Look for stable hotspots across repeated runs and ask whether the samples correspond to the scenario that matters. A short run may have too few samples; a high-frequency interval can perturb scheduling and increase overhead. A long run that includes unrelated startup, idle time, and shutdown dilutes a focused workload. Use the same warm-up, scenario, duration, and system load for before/after comparisons.
The tool may only attribute samples to functions when it can map addresses to loaded images and symbols. If output is mostly unknown addresses or module names, check whether the binary has useful symbol information and whether the image was built for the tested revision. Do not guess function names from an address without the matching image and symbols.
Use Callgrind output when a call graph helps
The -v <directory> option asks profile to write Valgrind/Callgrind-style output. Current source also enables full-stack analysis for that mode and uses a deeper caller stack, which changes both output detail and cost. Choose a new output directory under a user-writable location, keep the complete result, and open it with a compatible viewer such as QCachegrind if installed.
Example command shape:
profile -v /boot/home/Desktop/profile-run-01 -i 1000 /path/to/application
The profiler creates the output directory and per-entity files when the run finishes. Do not reuse a directory containing a previous profile unless the installed tool’s behavior is understood; separate runs make comparisons auditable. Callgrind-style visualization helps inspect call relationships but does not convert statistical samples into exact execution traces.
The basic -o <file> option writes the text results to a file. Keep error output as well as the result file, because launch failures, missing symbols, or output permissions can otherwise look like an empty profile. A file’s existence is not proof that a representative workload was captured.
Tune options conservatively
-f tells the tool to analyze the full caller stack and increases the default stack depth to 64. This can reveal more call relationships, but it consumes more work and can make output larger. Use it when the call chain is the question, not as an automatic setting for every run.
-s <depth> controls the number of return-address samples from the caller stack per tick. A deeper stack can help attribute indirect work, but it increases overhead and can expose less useful frames if symbols are incomplete. -c and -C change whether child threads or child teams are profiled, so changing them between runs changes the scope of the experiment. Record every option in the report.
Do not treat -a system-wide results as directly comparable with a command-mode run. The profiler may observe many teams and background activity. Narrow by a controlled system state or run the application directly when possible.
Turn a hotspot into a tested change
Use the profile to form one specific hypothesis: a function may perform too many allocations, repeatedly parse the same data, or redraw a large surface. Then inspect the relevant code, change one behavior, and rerun the same scenario with identical options. A lower sample count can mean less CPU use, but it could also mean the workload did less work or produced incorrect output. Verify functional results and elapsed time as well.
For UI stalls, pair CPU profiling with thread-state or I/O evidence. If the main thread is blocked, the sampled hotspot may be in a worker, or there may be no CPU hotspot at all. A profile cannot prove why a thread is waiting. Use the native Debugger, application logs, network/system tools, or kernel diagnostics according to the failure layer.
Do not publish a profile directory without reviewing it. Output can contain local executable paths, usernames, module names, and application activity. Retain the minimum files needed for reproducibility and redact private information before sharing.
Acceptance and report checklist
Accept a performance claim only when the same workload, binary, settings, and profile options were measured before and after, the output identifies a repeatable hotspot or reduction, the application still produces correct results, and elapsed time or responsiveness improved. Report sample interval in microseconds, run duration or workload completion, whether child threads/teams were included, whether kernel frames were sampled, and whether Callgrind output was used.
If two runs disagree, do not average away a large difference without investigating. Compare warm-up, CPU load, background tasks, image symbols, and event sequence. Sampling uncertainty is expected; repeat runs and report variability honestly.
Keep each result tied to the exact executable and workload revision. A profile from one binary cannot safely name functions in a different build just because the command and source tree look similar; addresses and inlining can change. Record a build identifier or source revision with the output, and preserve the original text result alongside any converted visualization. If the application is optimized, note the build configuration because optimization changes symbol availability and code shape.
Use a simple comparison table for repeated trials: run number, elapsed workload time, sample interval, total reported samples, top observed functions, and relevant system load. Compare like with like and retain outliers for investigation rather than deleting them without explanation. A profile is a clue for where CPU was observed, not proof of a causal performance defect. The final acceptance test should still exercise the user-visible behavior and confirm that the suspected optimization did not merely skip work.
Haiku’s profile is most useful when treated as a low-overhead measurement aid with a controlled question. It samples execution rather than tracing every call, its interval is in microseconds, and its scope changes with options. Pair it with a reproducible scenario, symbol-aware interpretation, and an independent functional/performance check before claiming an optimization.
Related:
- How to Debug a Native Haiku Application with Debugger
- How to Monitor System Activity with ProcessController on Haiku
Sources: