Skip to content
LinuxDeep Dive Published Updated 5 min readViews unavailable

Linux Zoned Storage: Write Pointers, Zone Reports, and zonefs

Operate Linux zoned block devices by checking zone models and write pointers, preserving sequential-write rules, and validating zonefs constraints.

Zoned block devices divide their address space into zones with different write rules. A sequential zone accepts writes only at its current write pointer and advances that pointer as data is written; reset operations return a zone to its beginning. This model appears in host-managed or host-aware disks and in zoned namespaces, but the exact device contract depends on its interface and model.

The kernel does not make every zoned device behave like an ordinary random-write disk. Applications and filesystems need to preserve zone constraints or use a layer designed to manage them. zonefs exposes zones as files but deliberately keeps sequential-write semantics visible. It is not a general-purpose filesystem that transparently turns random overwrites into log-structured updates.

Identify the device model and geometry

Begin with read-only attributes and the device’s reported topology:

lsblk -o NAME,TYPE,SIZE,ZONED,MODEL
cat /sys/block/sdX/queue/zoned
blkzone report /dev/sdX

Replace sdX with the actual whole block device and confirm the path before using any tool. Device attributes and util-linux options vary by kernel and package version. The zoned sysfs attribute reports none, host-aware, or host-managed when supported. A drive-managed device may appear as an ordinary block device because it does not expose host-visible zone commands.

The report should include zone start, length, capacity, type, condition, and write pointer where the device provides them. Zone length and usable capacity are not always equal. Conventional zones permit ordinary access; sequential zones have a write pointer and state transitions. Do not derive a write location from file length alone if the device reports a separate capacity.

Confirm that the complete storage stack preserves zoned information. RAID, device-mapper, virtualization, USB bridges, and enclosures may hide or transform capabilities. A zoned leaf device below an ordinary virtual disk does not guarantee the application sees a usable zoned interface.

Sequential writes are an application contract

For a sequential zone, the next write must begin at the write pointer and respect alignment, maximum append size, and zone capacity. A failed or partial write can leave the pointer at a different position than the application expected. On restart, query the device state rather than trusting only a userspace journal.

A zone append operation lets the device or block layer choose the actual append location and report it back, which can simplify concurrent writers when supported. It does not allow arbitrary overwrites. The application must record the returned location and ensure each zone has a defined owner or concurrency protocol.

Zone reset is destructive for the contents of that zone. It should occur only after the application has retired the data and updated its mapping or log. Do not reset a zone to recover from an unknown error without reconciling application metadata; that can make live records unreachable.

zonefs intentionally exposes the constraints

zonefs presents zones as files grouped by type, with little on-disk metadata. A sequential-zone file is append-oriented; direct I/O is used for writes, and write order must obey the device pointer. Conventional zones can be exposed differently, including aggregation depending on formatting options. The kernel documentation explains that zonefs aims to simplify application access while keeping the underlying zoned model explicit.

Formatting with mkzonefs writes filesystem metadata and should be treated as destructive. Validate the device identity, zoned model, intended capacity, and backup state before formatting. On a lab device, inspect the resulting directory tree and verify which files are conventional and sequential. Do not run formatting commands against a production device as an exploratory step.

The application needs a zone allocator, per-zone write position, recovery log, and policy for full, offline, read-only, or open zones. It must decide how to handle a failed write, restart after crash, and reclaim data. zonefs does not provide a database transaction manager or garbage collector for application records.

Observe zone state without disturbing data

Use zone reporting and sysfs queue attributes to inspect state, and correlate with block-layer errors. Keep the exact device path, kernel, firmware, zone size and capacity, open-zone limits, write-pointer state, and application mapping. A mismatch between the application journal and device write pointer is a recovery issue; do not hide it by resetting the zone.

The block layer has rules for scheduling writes to sequential zones. A scheduler or stacked target that reorders requests incorrectly can violate sequential delivery. Test the whole path, including device-mapper or RAID layers, rather than a direct-device benchmark alone. Confirm that the target supports the required write operation and reports the final sector correctly.

Host-managed and host-aware labels describe different levels of host responsibility. Read the device documentation and kernel attributes instead of assuming that the word zoned guarantees identical enforcement. Device firmware, transport, and controller behavior all affect the path from an application append to durable media.

Capacity, open zones, and write amplification

Zone geometry affects usable capacity and concurrency. A device can limit the number of open or active zones, and an application that opens too many may receive resource errors even with free bytes available. Keep a bounded working set of zones, close or finish them according to device policy, and monitor per-zone state.

Sequential allocation may reduce internal write amplification for certain media, but applications still perform cleaning or garbage collection when old data is superseded. If the workload has random updates, it must create a log-structured layer or use a filesystem that handles zoned constraints. Moving the complexity into userspace does not make it disappear; it makes recovery and testing part of the application.

Benchmark with the intended record size, concurrency, zone allocation, reset cadence, and recovery process. Throughput from a sequential write test says little about random updates or zone cleaning. Include device reset, power loss, and near-full conditions with disposable test data.

Acceptance criteria

Before production, verify that each layer reports the expected zoned model, the application never writes behind the pointer, zone reset follows durable metadata updates, and recovery reconciles device and application state. Test full zones, open-zone pressure, short writes, device removal, and reboot. Keep a recovery tool that can report state without issuing writes.

Zoned storage is reliable when the write pointer and data log are treated as shared state across the application, block layer, and device. The right workflow discovers the model, enforces sequential ownership, and treats reset as a destructive state transition rather than a generic retry.

Related:

Sources:

Comments