FreeBSD ZFS Snapshot Holds: Protect and Release Recovery Points
Use named ZFS snapshot holds to protect recovery points from deletion, inventory references, plan releases, and avoid confusing retention with backup.
ZFS user holds are named references that prevent a snapshot from being destroyed while the hold exists. They are useful when a backup job, replication workflow, incident response process, or human operator needs a snapshot to remain available beyond the normal retention schedule. A hold is a deletion guard, not an extra snapshot, replica, backup, or guarantee that the pool will remain healthy.
The lifecycle is simple to state and easy to get wrong operationally: create or identify the snapshot, apply a unique tag, inventory the hold, transfer ownership deliberately, and release the same tag only after the dependent workflow has finished. Because a held snapshot refuses ordinary destruction with EBUSY, orphaned holds can silently retain snapshots and pool space. Every hold should have an owner, purpose, and release condition.
Inspect before applying a hold
Confirm the dataset, snapshot name, pool health, and available space before creating new retention state:
zpool status -x
zfs list -t snapshot -o name,creation,used,refer -r tank/data
zfs holds -r tank/data@daily-2026-10-03
The snapshot name is only an identifier; confirm that it represents the data and time required by the recovery plan. zfs holds lists user references on a snapshot. -r includes descendant snapshots where applicable. The exact zfs list properties available can vary by version, so use zfs get or the installed zfsprops(7) manual if a field is unknown.
A snapshot is not independent of its pool. If the pool or its only devices are lost, a hold does not restore the data. Check that the dataset is part of the intended recovery scope, that backups or replication are current, and that a restore has been tested. A held source snapshot on the same pool is still subject to pool failure, administrative destruction of the pool, and other catastrophic events.
Apply an owner-specific tag
Use a tag that identifies the consumer and change, not a generic word such as keep:
zfs hold -r backup-job-20261003 tank/data@daily-2026-10-03
zfs holds -r tank/data@daily-2026-10-03
The -r option applies the hold recursively to descendant file-system snapshots. Each snapshot has its own tag namespace, and tags must be unique within that snapshot. A tag such as backup-job-20261003 creates a discoverable link between the hold and its owner. If several independent workflows protect the same snapshot, they should use distinct tags so one process cannot accidentally release another process’s reference.
The command is idempotent only in the broader orchestration sense if your automation handles the “already exists” result deliberately. Do not blindly retry a failed hold and interpret any nonzero result as success. Query the exact snapshot’s holds, check the job identifier, and record which snapshots were changed. Avoid recursive holds across an unbounded hierarchy unless the retention policy intentionally protects all descendants.
Inventory holds before deletion or retention cleanup
Before a retention job destroys old snapshots, list holds in a format suitable for review:
zfs holds -r -H -p tank/data@daily-2026-10-03
zfs list -t snapshot -o name,creation,used,refer -r tank/data
The -H option omits headers for tabular processing, while -p prints timestamps as Unix epoch values in versions documenting those flags. Verify flags for the installed release. If a scheduled snapshot cleanup fails with EBUSY, treat that as evidence of an active hold. Do not add force flags or destroy a parent dataset to get around it. Locate the tag owner and confirm the downstream workflow has finished.
Track holds in a small inventory containing pool, dataset, snapshot, tag, creator, purpose, creation time, dependent job, owner, expiry condition, and release approval. The ZFS hold itself is not a complete business metadata record; the tag string and list output may not tell you why it was applied. External inventory is needed for aging and accountability.
Holds can increase retained data indirectly because they prevent deletion of the snapshot and the snapshot continues referencing blocks that would otherwise become eligible for reclamation. They do not make a second copy or directly reserve an exact amount of space. Actual space impact depends on blocks shared with live datasets and other snapshots. Measure with ZFS properties and pool capacity; do not estimate retained bytes by adding snapshot used values without understanding sharing.
Release precisely and test the lifecycle
Remove a hold only when its named dependency is complete:
zfs holds -r tank/data@daily-2026-10-03
zfs release -r backup-job-20261003 tank/data@daily-2026-10-03
zfs holds -r tank/data@daily-2026-10-03
The release tag must exist for each targeted snapshot. Recursive release targets descendants, so inspect the scope before using -r. A typo or an overbroad dataset argument can fail partway or release more references than intended. Query before and after, retain command output, and have the data owner approve releases that affect recovery points.
Do not release all holds as a generic cleanup. If a snapshot has multiple tags, each can represent a different owner. Removing one reference must not be interpreted as permission to delete the snapshot while other holds remain. Conversely, a snapshot with no user hold may still be required by a retention policy or incremental replication chain; the absence of a hold does not make it safe to destroy.
Test the policy on a disposable dataset before applying it to production. Create a test snapshot, apply a tag, confirm that ordinary destruction is blocked, release the tag, and confirm that the reference is gone. Do not run destructive tests against production data. A test that proves ZFS hold semantics does not prove a backup application will create or release the intended tags.
Coordinate holds with automation and replication
An automatic snapshot tool and a hold solve different problems. A schedule creates points in time and a retention rule prunes them. A hold protects a selected point from that pruning until an external condition is met. If the tool does not understand holds, cleanup may log failures and retain old snapshots; alert on that condition and assign an owner rather than disabling cleanup.
Replication adds another lifecycle boundary. A hold on the source does not itself prove that a receiver has a complete copy, and a successful send does not automatically mean the source may be released. Confirm the receiver’s snapshot, stream completion, integrity checks, and restore procedure before removing a protection tag. If a workflow needs the snapshot to remain available for incremental send, record the exact dependency and release it only according to the replication design.
Do not build a cleanup script that releases tags based solely on age. Age does not establish that a backup finished, an incident is closed, or a legal hold ended. Use a state machine or explicit completion signal from the consumer, validate the exact snapshot and tag, and log each release. Run cleanup in dry-run or report-only mode first where the surrounding automation permits.
Incident response and failure cases
When a snapshot cannot be destroyed, inspect the exact snapshot and its user references, then identify the tag’s owner. When the owner is unknown, retain the snapshot while investigating; space pressure is a reason to plan capacity, not to erase unexplained recovery state. Escalate according to data ownership policy and capture pool status and snapshot properties before and after.
When a recursive hold or release affects fewer descendants than expected, verify the dataset tree and snapshot naming. A recursive operation applies to the relevant descendant snapshot set, not to snapshots that were never created. Check for spelling mistakes, different snapshot timestamps, and dataset boundaries. Never assume a similarly named snapshot exists on every child dataset.
Acceptance means every protected snapshot has a unique documented tag and named owner, routine retention reports but does not delete protected recovery points, release occurs only after an independently verified dependency completion, and a restore or replication test confirms the recovery point is useful. A hold is a narrow deletion guard inside a broader data-protection design.
Related:
- How to Automate ZFS Snapshots with periodic
- How to Replicate FreeBSD ZFS Datasets Safely with send and receive
Sources: