Skip to content
SRE & DevOpsDeep Dive Published Updated 9 min readViews unavailable

etcd Compaction and Defragmentation: Control MVCC Growth Safely

Separate etcd history compaction from backend defragmentation, select retention deliberately, monitor quota headroom, and recover NOSPACE safely.

An etcd database can keep growing even after an operator has deleted old objects, because etcd retains a history of key revisions and its backend file can contain reusable but not yet returned space. Two maintenance operations address different layers: MVCC history compaction discards revisions older than a chosen point, while backend defragmentation rebuilds a member’s database file to return free pages to the host filesystem. Running only one may not achieve the result an incident responder expects.

This distinction is operationally important for Kubernetes control planes. Compaction affects how far back a client can watch or read; defragmentation blocks reads and writes on the member being rebuilt. Both must be planned around client behavior, quorum health, disk headroom, backups, and the exact etcd version. Managed control planes often own these operations and should be maintained through the provider’s supported process, not direct access to their database.

Three different meanings of “compaction”

The word is overloaded in etcd operations:

  1. MVCC revision compaction removes historical key-value versions before a revision. This is the history-retention operation described in this guide.
  2. Raft log snapshotting and truncation controls replicated consensus-log entries and follower catch-up behavior. It is not a substitute for MVCC history retention.
  3. Backend defragmentation rebuilds the local backend file after old revisions have been compacted. It releases internal free space back to the filesystem, but temporarily blocks reads and writes to that member.

Deleting a Kubernetes object creates newer keyspace state; it does not automatically make every prior revision immediately disappear from MVCC history or shrink the backend file. Likewise, an etcd snapshot backup is a separate recovery artifact, not a compaction mechanism.

What MVCC history compaction changes

etcd assigns revisions to changes in its keyspace and retains old versions for a window of history. Compaction at revision R makes older revisions unavailable. A client that asks for a read or watch starting before the compacted boundary can receive a compacted-revision error and must establish a fresh view of current state. Well-behaved Kubernetes controllers and clients list again and start a new watch; custom clients must implement the equivalent relist-and-resynchronize path rather than retrying the same expired revision forever.

The retention window therefore has a correctness and availability dimension, not just a disk dimension. It must exceed the realistic time a disconnected or overloaded client may need to catch up. At the same time, retaining revisions indefinitely consumes backend capacity. Choose the retention model from actual write volume and client behavior, then test slow watchers and reconnects in a non-production environment.

etcd supports time-based periodic compaction and revision-based compaction. Periodic retention is easier to reason about in elapsed time: for example, the etcd v3.7 maintenance guide describes a 10h history window as the general-purpose default, while use cases with frequent updates or infrequent updates may choose shorter or longer windows. Revision retention instead keeps an approximate count of recent revisions; its elapsed-time window can vary dramatically with write rate. A quiet cluster and a busy control plane do not get the same hours of history from the same revision count.

The example below shows the form of a time-based setting, not a universal recommendation. For Kubernetes, use the version- and distribution-supported etcd configuration mechanism, and verify whether the provider already manages automatic compaction:

--auto-compaction-mode=periodic
--auto-compaction-retention=10h

Manual etcdctl compact REVISION can be useful in a carefully planned maintenance operation, but it is easy to choose a boundary too far ahead of a lagging consumer’s last observed revision. Record the current revision and the approved retention boundary, understand which clients depend on historical watches, and confirm their relist behavior before compacting. Do not use an arbitrary “latest revision” compaction as a routine disk cleanup shortcut.

Why compaction may not shrink the file

Compaction makes obsolete MVCC history eligible for removal from the live keyspace. The backend may then have internal free pages that etcd can reuse for future writes. Those pages are not necessarily returned to the operating system, so df can show little or no change even though logical space has been reclaimed.

Defragmentation rebuilds a member’s backend and returns reusable free space to the filesystem. It is a per-member operation; it is not replicated to the other members. While a live member is being defragmented, that member cannot serve reads or writes normally. The cluster can continue only if the remaining healthy members retain quorum and client traffic can tolerate the reduced capacity. In a three-member cluster, defragment one member at a time and wait for it to rejoin and catch up before touching another. Never launch concurrent defragmentation across the voting quorum.

Compaction usually comes before defragmentation: defragmenting a database that still contains the old revisions does not remove that history. Repeated defrag without new compaction may simply rebuild a nearly equally large backend. Conversely, compaction alone may not return disk blocks to the host filesystem. Use measurements to decide which layer needs action.

Read the right size metrics

The useful comparison is between logical space in use, physical backend size, quota, and filesystem capacity; one chart cannot show the whole picture. The etcd metrics documentation exposes separate measurements for backend bytes in use and total physical database bytes. Their names have changed across etcd releases, so inspect the actual /metrics exposition and match it to the deployed version instead of copying a metric name from an old dashboard.

When physical size is much larger than in-use size, compaction may already have created reusable space and defragmentation may recover filesystem blocks. When both values are growing together, history compaction or excessive object churn deserves attention. When free filesystem space is low but the etcd backend is below its quota, the host can still fail writes because the filesystem, not the etcd quota, is the immediate limit. Monitor all three layers and alert with enough headroom for the time required to take action.

For a first-pass endpoint health and status check, etcd v3.7 documents commands of this general form:

etcdctl endpoint status --cluster --write-out=table
etcdctl endpoint health --cluster
etcdctl alarm list

These examples omit TLS flags and credentials. In production, use the authenticated endpoints and certificates for the intended cluster, confirm every expected member appears, and compare member health and backend size. The --cluster flag discovers members from the cluster; do not assume a command aimed at one endpoint has checked every member.

A safe maintenance sequence

Use an explicit change window for self-managed etcd, especially before defragmentation. A disciplined sequence reduces the risk of converting a storage-maintenance task into a quorum or recovery incident:

  1. Identify ownership and version. Confirm whether the control plane is self-managed, which etcd version and etcdctl build are running, how TLS is configured, and which system owns automatic compaction. Managed Kubernetes control planes should use provider procedures.
  2. Establish recoverability. Take and verify a recent etcd snapshot using the release-compatible tooling. Store it away from the member disks and test the restore procedure on an isolated cluster. Compaction is intentionally destructive to old revisions; it is not reversible by changing the retention flag afterward.
  3. Capture a baseline. Check health and status for every member, leader and quorum state, current and compact revisions, backend size and size-in-use metrics, free filesystem space, disk latency, and active alarms. Resolve unhealthy members before maintenance.
  4. Choose retention. Set an approved time or revision window based on watch-client recovery expectations and write rate. Roll out the same supported configuration to each control-plane member as required by the distribution, and verify what the running servers actually loaded.
  5. Compact history. Prefer the documented automatic policy for steady-state retention. If manual compaction is necessary, use the agreed revision boundary and communicate that clients older than it must relist. Monitor API-server and controller error rates for compacted-watch errors.
  6. Defragment sequentially. Recheck quorum before each member. Defragment only one member, monitor its availability and latency, wait for it to return and catch up, then verify the cluster is healthy before proceeding to the next member. Exact orchestration differs by etcd release and vendor tooling; follow the supported procedure.
  7. Verify the result. Compare in-use and physical backend sizes, filesystem free space, endpoint health, alarm state, API-server write success, and controller watch recovery. Record measurements and next maintenance criteria rather than declaring success from a completed command alone.

Do not confuse this procedure with changing --snapshot-count. That setting controls Raft log retention and follower catch-up trade-offs, not the MVCC history window used by key-value watches.

Responding to a NOSPACE alarm

When a member crosses the backend quota, etcd can raise a cluster-wide NOSPACE alarm and restrict normal operation to reads and deletes. This is a protective mode, not evidence that every failed client write was absent. The etcd maintenance guide documents that a request can return a no-space error while the backend apply path still records the transaction and raises the alarm. Before retrying a business operation, check its postcondition or idempotency key; blind retries can duplicate external side effects.

The recovery order is deliberate: identify the owning keys and approved deletion/retention action, compact enough obsolete history to reduce logical use, defragment each member so physical backend files fit, and only then clear the alarm using the procedure for the deployed release. Clearing the alarm before all relevant members have safe headroom can immediately re-trigger it. Never delete arbitrary Kubernetes keys from etcd to “make room”; use the Kubernetes API or the owning system’s supported cleanup path so controllers and finalizers can maintain consistency.

If the alarm blocks API-server writes, prioritize service impact assessment and communicate that cluster changes may be rejected. Preserve logs, metrics, and a verified snapshot. For provider-managed etcd, contact the platform’s support or follow its documented recovery workflow rather than improvising member commands against an opaque control plane.

Operational acceptance criteria

  • Each etcd member is healthy, has a leader and quorum, and is within storage and filesystem budgets.
  • The chosen retention window matches client watch recovery needs and is consistent across the supported topology.
  • Monitoring distinguishes logical backend use, physical backend size, quota, and filesystem free space.
  • Defragmentation is member-by-member with a health and catch-up gate between members.
  • Snapshots are verified independently; maintenance success is not treated as backup success.
  • Compacted-watch recovery is tested for custom clients, and applications handle ambiguous no-space write results idempotently.
  • Alarm clearing occurs only after storage headroom and endpoint health are verified for the full cluster.

Compaction, defragmentation, and snapshots solve different problems: retention, filesystem reclamation, and recovery. Production maintenance is safe when the operator knows which problem the measurements show, preserves enough history for clients, protects quorum during per-member work, and verifies the application-facing API after every change.

Related:

Sources:

Comments