FreeBSD zpool iostat: Measure ZFS Vdev Activity Without Misreading Samples
Interpret FreeBSD zpool iostat counters, first reports, interval samples, vdev breakdowns, and the boundary between logical and physical I/O.
zpool iostat reports logical I/O statistics for ZFS pools and their virtual devices. It is useful for finding uneven vdev activity, observing throughput over time, and correlating storage behavior with workload changes. It is not a direct measurement of physical disk service time: ZFS can aggregate nearby writes, issue extra I/O for redundancy, and perform work that is not represented as a one-to-one application operation.
The most common interpretation error is treating the first report as a short interval sample. On the FreeBSD zpool-iostat(8) interface, the first report is normally statistics accumulated since boot. With an interval and count, later reports represent repeated observations; -y suppresses the since-boot first report. Read the installed manual because options and column sets can vary across bundled OpenZFS versions.
Start with a topology and a quiet baseline
Before attributing activity to a disk, record the pool’s vdev tree and health:
zpool status -v tank
zpool list -v tank
zpool iostat -v tank
The pool and vdev hierarchy matters. A mirror, RAID-Z group, special vdev, log device, and cache device serve different roles. A child vdev with higher activity is not automatically defective or imbalanced; it may reflect the configured topology, workload locality, or maintenance such as a scrub or resilver. Make the mapping between device paths and vdev roles explicit before acting.
Record the time, workload, and pool state. A single snapshot can show accumulated counts but cannot establish a trend. Compare observations over a defined interval during a quiet baseline and during the workload of interest. Do not change pool layout or tuning based on one iostat screen.
Sample interval statistics deliberately
Use a bounded interval and count when gathering a repeatable sample:
zpool iostat -v -y -T d tank 5 12
This asks for verbose per-vdev output, suppresses the since-boot first line, emits a date-format timestamp, and requests 12 reports at five-second intervals. Check that the installed release supports each option before using it in automation. A long-running monitor can omit the count, but an operational capture should have an explicit end condition.
Use -p when a parser needs raw numeric values rather than human-scaled units. Preserve the original output and command line, because formatting and columns can differ across versions. Avoid parsing by fixed column position without checking the installed manual and the actual header emitted by that host.
With periodic reports, compare the same vdev across successive samples. A pool-wide total can hide one slow member; a per-device report can reveal skew, but skew alone does not establish the cause. A burst of writes may be routed unevenly by design, and redundancy levels can amplify physical work relative to application writes.
Separate pool statistics from device statistics
The zpool iostat manual describes logical pool and vdev I/O statistics. iostat(1) can provide physical device observations. Use both when diagnosing, but do not add their byte counts together as if they measured independent work. They observe different layers of the storage stack, and the same operation may appear in each with different aggregation, timing, or queue behavior.
Correlate the sample with zpool status, gstat, device health reports, and workload metrics. If a vdev is busy while physical disk counters are quiet, consider caching, aggregation, device mapping, or sampling-window mismatch. If physical activity is high but pool-level activity is low, look for unrelated swap, logs, another pool, or block-device users outside ZFS.
Diagnose common shapes without jumping to conclusions
High read throughput across all vdevs during a backup may be expected. Uneven read distribution during a random workload may reflect dataset layout or cache effects, not necessarily a bad disk. One member with markedly higher latency or queue pressure is a lead for investigation, not proof of failure. Check the device health, controller logs, zpool status -v, and whether scrubs or resilvers are active.
A pool with low throughput can still have high latency or be blocked on a slow member. Conversely, high throughput does not prove acceptable application latency. The version of zpool iostat may expose queue or latency options such as histograms on supported OpenZFS releases; use those only after confirming the installed syntax and understanding what the fields include. Do not mix queue time with device service time.
When the first output line is unexpectedly large, check whether -y was omitted. That line is useful as a lifetime or since-boot view, but it is not a five-second rate. For automation, label the report type and interval explicitly so dashboards do not treat the cumulative line as a delta.
Build an evidence capture for an incident
Capture topology, status, samples, and hardware data with timestamps:
date -u
zpool status -v tank
zpool iostat -v -y -T d tank 2 30
iostat -x 2 30
The exact iostat(1) flags should be checked on the target release; physical-device options are not identical on every Unix. Preserve command output and host release. If the issue affects only one workload, run a small controlled reproduction and record the application-level latency at the same time as storage samples.
Avoid invoking scripts through zpool iostat -c without reviewing their trust and privilege model. Some versions can execute vdev scripts to add columns, and the manual describes conditions for privileged script execution. A diagnostic option that runs code is not merely a formatting switch. Use scripts from a controlled path and review their inputs and permissions.
Turn observations into operational checks
Use zpool iostat to answer bounded questions: which vdev is receiving I/O, whether a maintenance operation is changing the pattern, and whether activity persists across a defined interval. Pair it with a known pool topology, disk health, and workload-level latency. Do not use it as a standalone SMART replacement, a data-integrity check, or a reason to replace hardware.
A useful acceptance record includes the FreeBSD release, OpenZFS version if available, exact command, report timestamps, pool status, workload phase, and any active scrub or resilver. For a suspected imbalance, compare repeated samples under equivalent workload and account for cache state and vdev type. For a suspected failure, correlate with checksums, device errors, controller events, and the pool’s redundancy margin.
The goal is an evidence chain from application symptoms to ZFS vdev activity to physical device behavior. Each layer answers a different question. Keeping those boundaries explicit prevents a visually dramatic counter from turning into an unsupported hardware diagnosis.
Make comparisons reproducible
If a baseline is used for capacity or incident review, store the raw output with a pool identifier and the system’s release. Include whether the sample began during boot, a scrub, a resilver, a snapshot send, or a scheduled backup. A change in the first-report behavior can otherwise look like a sudden throughput change when only the sampling window changed.
When comparing two pools, normalize the question rather than comparing headline totals. Record vdev count and type, device model, workload, interval length, and whether the run was warm or cold from cache. A four-device mirror and a four-device RAID-Z group have different data and parity behavior; similar total bytes do not imply similar physical work or performance expectations.
If a monitoring system consumes output, test it against the exact FreeBSD and OpenZFS versions deployed. Require a header check, reject missing pool rows, and track the sample timestamp. Alert on sustained patterns combined with application impact, not on a single counter crossing a generic threshold. This keeps collection failures, maintenance traffic, and genuine device trouble from becoming indistinguishable alerts.
Related:
- FreeBSD GEOM I/O Observability: Read gstat and iostat Without Double Counting
- FreeBSD systat Operations: Read Live CPU, Memory, Network, and I/O Views
Sources: