Skip to content
SRE & DevOpsDeep Dive Published Updated 9 min readViews unavailable

Prometheus TSDB Capacity and Retention: WAL, Blocks, and Backups

Size and operate Prometheus local storage using WAL and block lifecycle, current retention configuration, compaction headroom, and consistent backups.

Prometheus local storage is a database lifecycle, not a directory that can be capped safely by setting a PVC to an arbitrary size. The TSDB keeps recently ingested samples in a mutable head protected by a write-ahead log (WAL), persists older data into time-based blocks, compacts those blocks in the background, and removes old blocks according to retention. During compaction, source and replacement blocks can coexist, while WAL and head data consume disk beyond the persistent-block retention budget.

An operator who sizes only for the retention window can hit a full filesystem before Prometheus evicts old blocks. An operator who copies the live data directory without a consistent snapshot can restore a gap in the newest samples. This guide explains the storage lifecycle, current configuration model, practical capacity planning, backup boundaries, and safe diagnosis.

Understand the TSDB directory before sizing it

Prometheus stores ingested samples in two-hour blocks. Each block contains sample chunks, an index, and metadata. The active head block is mutable and remains in memory while its recent data is protected on disk by WAL segments. Prometheus replays the WAL on startup to reconstruct data that had not yet become a completed block. Deleting WAL files to save space can therefore discard recent samples or prevent a clean recovery.

Completed blocks are compacted into larger time ranges in the background. During compaction, the original blocks and the new compacted block must both fit on disk until cleanup removes the originals. Tombstones created by series deletion also do not immediately rewrite chunk files; physical reclamation happens later through compaction. A retention setting therefore controls the age or size of stored blocks, not every byte Prometheus can use at any instant.

The local TSDB is not itself clustered or replicated. A persistent volume protects data from an individual container restart, but it does not automatically protect against a node, volume, or zone failure. Prometheus local storage should use a filesystem supported for its database workload; upstream documentation warns that network filesystems such as NFS are not supported because they can cause unrecoverable corruption. If the design requires larger-scale durability or distributed querying, evaluate a remote storage architecture with its own documented failure and retention guarantees rather than mounting the same local TSDB directory from multiple replicas.

Configure retention in the current Prometheus configuration file

Current Prometheus configuration documentation exposes reloadable TSDB settings under storage.tsdb. For example:

storage:
  tsdb:
    retention:
      time: 30d
      size: 120GB

The values are illustrative. Set a retention window and size budget that match the product’s query and recovery needs and the actual volume allocation. Current Prometheus 3.x documentation marks the command-line retention flags as deprecated in favor of the configuration file fields. Existing deployments may still contain flags, but new configuration should follow the version’s current supported reference and be checked with the matching promtool binary before rollout.

When both time and size limits are set, whichever policy triggers first removes older blocks. A shorter time value can discard history even when the volume has space; a tighter size budget can shorten the effective history during an ingestion spike. If neither policy is explicitly set, Prometheus has a documented default retention time; set it intentionally rather than assuming that a larger PVC automatically means longer retention.

The size setting is a limit for persistent blocks, not a hard upper bound for the entire data directory. WAL and memory-mapped head chunks are counted in on-disk use but are not the data removed to satisfy block retention. Compaction can temporarily exceed the configured block limit. The current storage guide recommends sizing the retention limit to at most roughly 80-85% of the allocated Prometheus disk, keeping the remaining space as headroom for compaction. The exact safe margin depends on ingestion rate, series churn, block sizes, and workload behavior; monitor real peak usage instead of treating the percentage as a guarantee.

The current configuration reference also describes percentage-based retention as experimental. Do not adopt an experimental storage policy in production without validating its behavior against the exact Prometheus build and filesystem environment. For a well-understood baseline, use explicit time and/or size retention, capacity alerts, and a tested recovery process.

Estimate capacity from ingestion and measure it in production

A rough starting estimate is:

block_capacity ~= samples_per_second * retention_seconds * average_bytes_per_sample

Prometheus documentation gives a rough average of about 1-2 bytes per sample for local storage, but actual cost depends on label sets, churn, metric types, scrape patterns, and compression. The formula estimates sample blocks; it does not include WAL peaks, head chunks, filesystem metadata, compaction overlap, or safety margin. Use observed disk growth over representative high-ingestion periods to replace the estimate before committing to a PVC size.

A reliable capacity review should include:

  • Peak and steady-state ingested samples per second, including deployments that temporarily increase target or series counts.
  • Retention time needed for investigation, compliance, and operational querying, separated from backup-retention requirements.
  • Block bytes, WAL/checkpoint usage, head-chunk usage, and compaction peak during the same period.
  • The actual filesystem capacity available to Prometheus, not just a container request or PVC nominal size.
  • Growth scenarios for new instrumentation, additional scrape targets, higher cardinality, and remote-write buffering if configured.
  • Volume expansion and alert lead time sufficient to resize storage before the filesystem becomes full.

High-cardinality metrics can make storage growth much faster than sample-rate intuition suggests. A series is identified by its metric name and complete label set; a changing label value creates new series and additional index and metadata costs. Reduce unnecessary series at instrumentation or scrape relabeling rather than assuming that a longer retention flag will solve uncontrolled growth.

Alert on the Prometheus data filesystem with enough lead time for both response and compaction behavior. Track current free bytes and percentage, write rate, sample ingestion, WAL size, and head-series growth where those metrics are exposed by the deployed version. A filesystem alert alone is not enough if the time to exhaustion is shorter than a volume resize or retention cleanup can complete.

Back up the head and blocks consistently

Upstream recommends Prometheus TSDB snapshots for backups. Copying only completed block directories can lose data recorded since the latest block was created, while copying a live directory file by file can capture an inconsistent view of files being written or compacted. Use the supported snapshot mechanism or stop the process cleanly and follow a tested storage-level backup procedure. Restore tests matter: a snapshot that has never been opened by a matching server is not evidence that the recovery process works.

Prometheus exposes a TSDB snapshot API, but the admin API is disabled by default and should only be enabled and reachable through the deployment’s approved administrative access path. Configure authentication and network restrictions according to the environment, and disable the endpoint again if the operational design does not require it. Do not publish an unauthenticated management endpoint simply to make scheduled backups easier.

If a backup process intentionally excludes WAL and head directories, it produces a coherent but older view and loses the time range held only in those directories. Document that recovery point explicitly. If the backup includes them, use a snapshot or a filesystem procedure that guarantees consistency. Test that the backup contains the expected block range, that the target can restore it, and that remote-write or other integrations do not create confusing duplicate history during a recovery.

Keep backup retention separate from TSDB retention. The local server may remove blocks to meet its current policy; a backup system may preserve older restore points for a different period. Conversely, configuring a long TSDB retention does not protect against volume loss or operator deletion.

Investigate disk pressure without deleting database files

When disk use rises unexpectedly, first establish which part of the data directory is growing and what changed at the same time. Review Prometheus logs for compaction, WAL replay, retention, and filesystem errors. Compare ingestion and series counts to the last healthy period, inspect scrape target and label changes, and check whether an in-progress compaction is temporarily holding both old and new blocks.

In a Kubernetes deployment, inspect the Prometheus Pod’s PVC, the storage class and backend expansion behavior, and the node filesystem if local PVs are used. Confirm the container’s actual mount path and filesystem usage instead of assuming that a larger PVC was mounted or expanded. A size limit that is lower than the usable block budget can cause retention to delete data earlier than expected; a very high limit can leave the volume with no headroom before block cleanup is eligible.

Avoid removing arbitrary block or WAL directories as routine cleanup. That can erase a time range, break recovery, or leave state inconsistent. If a TSDB is corrupted and will not start, preserve a copy of the data directory first, then follow the version-specific recovery guidance and identify which block or WAL range is damaged. Deleting files should be the last-resort recovery action with an explicit accepted data-loss window, not a substitute for retention planning.

Before changing retention, calculate the effect of both policies. A decrease in time can drop queryable history when the next block cleanup runs. A decrease in size can remove the oldest blocks sooner than the time policy. Increasing retention without increasing capacity can cause filesystem exhaustion. Validate the new config, reload or roll out according to the deployment’s configuration model, and observe the actual block cleanup cycle before declaring the issue resolved.

Production acceptance checklist

  • TSDB retention is configured in the version-appropriate configuration file, with time and size behavior understood.
  • The block budget leaves measured head, WAL, filesystem, and compaction headroom on the actual data volume.
  • Storage is persistent where required and uses a filesystem supported by the Prometheus TSDB.
  • Disk-free, ingestion, series-growth, and WAL signals have actionable alert thresholds before exhaustion.
  • Backups use a consistent snapshot or documented offline method and have been restored in a test.
  • Backup retention is independent from local TSDB retention, and the recovery point is documented.
  • Admin snapshot access remains within the approved management boundary and is not exposed as an unauthenticated public endpoint.
  • Runbooks prohibit ad-hoc deletion of WAL or block directories except under a data-loss-approved recovery procedure.

Prometheus retention is a policy for block history, not a quota for total filesystem use or a backup strategy. Plan for head and WAL growth, allow compaction to overlap source and output blocks, monitor the real data volume, and verify recovery with a consistent snapshot. That operating model makes retention predictable without sacrificing the most recent data to an emergency cleanup.

Related:

Sources:

Comments