Skip to content
SRE & DevOpsDeep Dive Published Updated 10 min readViews unavailable

Kubernetes LIST and WATCH: Consistent Caches, Recovery, and Scale

Build reliable Kubernetes API caches with LIST/WATCH resource versions, pagination, bookmarks, streaming initial state, and safe recovery from 410 Gone.

Kubernetes controllers and integrations often need a local view of many API objects and timely notification when that view changes. The API’s LIST and WATCH operations provide that pattern, but correctness depends on more than opening a long-lived HTTP stream. A client must establish an initial state, remember the collection’s checkpoint, process watch events without gaps, recover when history expires, and prevent duplicate or stale observations from corrupting downstream work.

The central rule is to treat the API as the authority and the client-side cache as a recoverable projection. Watches are not a durable event log, an exactly-once queue, or a promise that every intermediate object state will remain available. This article focuses on the Kubernetes API contract for collection reads, watch resource versions, pagination, bookmarks, and streaming initial events. Client libraries can implement much of the protocol; verify the exact behavior of the library and API server versions you deploy.

The LIST/WATCH handoff

A collection LIST returns both a set of objects and metadata.resourceVersion for the collection snapshot. A WATCH started from that collection resource version streams changes that occurred after the snapshot. This handoff closes the gap between “read current state” and “subscribe to changes”:

LIST collection
  -> populate local cache from items
  -> save ListMeta.resourceVersion
WATCH from that resourceVersion
  -> apply ADDED / MODIFIED / DELETED events
  -> save progress as watch events are processed

For example, if a Pod list response has metadata.resourceVersion: "10245", a watch request for that Pod collection can start with resourceVersion=10245. The string is a checkpoint associated with an API resource and server state, not a wall-clock timestamp or one universal cluster sequence. Preserve it exactly for the API client; do not compare a Pod checkpoint with a Deployment checkpoint or assume it fits a fixed-width integer. Current Kubernetes documentation defines additional orderability guarantees for conformant API servers, but client portability still benefits from treating the value as an API token unless a specific versioned use case requires comparison.

The resourceVersion on the list metadata is the collection checkpoint. An individual object’s metadata.resourceVersion says when that object instance was last modified; it is not the correct checkpoint for the entire collection. Starting a collection watch from an arbitrary item’s version can miss changes to other objects or provide semantics different from the snapshot the client actually loaded.

Reconcile state, not assumed event history

Watch events describe changes after a point, but clients should not build critical state transitions on the assumption that they will observe every intermediate mutation. A slow consumer, broken connection, API server restart, or expired history can interrupt a watch. Multiple updates may occur between observations, and a client may see only the latest object state when it resumes. This is why Kubernetes controllers use events to trigger reconciliation against current state rather than treating each event as an imperative command that must run exactly once.

For a cache consumer, make event handling idempotent. Upsert the full object for ADDED and MODIFIED; remove the key for DELETED; handle tombstones or deletion metadata according to the client library’s contract. Separate cache application from side effects such as provisioning cloud resources, sending notifications, or writing to another database. If an event is processed twice after reconnect or retry, the result should remain correct.

The BOOKMARK event is a progress marker, not an object mutation. It can advance the checkpoint to indicate that changes through a resource version have been delivered. A client may request bookmarks with allowWatchBookmarks=true, but the server chooses whether and when to send them. Never use an expected bookmark interval as a health guarantee, and do not discard pending events merely because a bookmark was requested.

Handle expiration as a cache rebuild

Kubernetes servers retain only a limited window of change history. If a client requests a watch from a version that is no longer available, the API can return HTTP 410 Gone. That is not a transient connection error to retry forever with the same version. The cached projection can no longer be proven complete from its checkpoint.

A safe recovery sequence is:

  1. Stop applying events from the expired watch and mark the local view unsynchronized.
  2. Clear or quarantine the cache so readers do not mistake it for a current complete snapshot.
  3. Perform a fresh LIST using the same resource type, namespace scope, and selectors.
  4. Atomically replace the old cache with the new snapshot and store the list metadata’s resourceVersion.
  5. Start a new WATCH from that checkpoint, then mark the cache synchronized once the initial snapshot and watch handoff are established.

If clearing the cache would cause harmful transient behavior, keep the previous snapshot explicitly marked stale and fence any decisions that require fresh state. Do not continue to serve it as authoritative. A bounded retry with backoff is appropriate for transient transport failures, but a 410 requires re-establishing state. Monitor relist frequency: repeated expiration can indicate slow event processing, high churn, reconnect instability, or expensive full LIST operations.

Paginate large collections without mixing snapshots

Large LIST responses can consume substantial API server and client memory. The limit and continue parameters allow a collection to be read in pages. The API’s continuation token preserves the snapshot and position for the next page; use the token exactly as returned and retain all other query parameters, including namespace and selectors.

GET /api/v1/pods?limit=500&labelSelector=app%3Dworker
  -> items, metadata.resourceVersion, metadata.continue

GET /api/v1/pods?limit=500&labelSelector=app%3Dworker&continue=<returned-token>
  -> next page from the same collection snapshot

Do not decode or synthesize a continuation token. It is opaque and valid only for the same list query apart from the token itself. A completed paginated list represents a consistent collection snapshot; changes after that snapshot are then obtained by watching from the collection resource version. This keeps a large initial load from silently mixing object states captured at unrelated times.

Continuation tokens expire. If the server returns 410 Gone for an expired token, restart the list from the beginning when a consistent snapshot is required. The API may provide a replacement continuation token that can continue from a newer snapshot, but blending that result with earlier pages creates an inconsistent view: objects could be omitted or represented at different points in time. If the application can tolerate that, make the choice explicit rather than silently treating the result as an atomic snapshot. Keep list pagination bounded and avoid scheduling expensive full relists for many clients at once.

When resourceVersion is supplied to a list request, also set resourceVersionMatch according to the documented semantics and server support. NotOlderThan can be more scalable than requiring the most recent quorum read when strong consistency is unnecessary. Older or non-conformant API servers may ignore the parameter, so check the response and test against every supported cluster version. Do not copy a resourceVersion from one resource type into a different type’s list and assume it means the same thing.

Streaming initial state

On large clusters, listing a huge collection and then opening a watch can be costly. Kubernetes supports sendInitialEvents=true on a watch request to stream synthetic ADDED events representing the initial state, followed by ordinary watch events. The API requires resourceVersionMatch=NotOlderThan with this option; the initial synchronization boundary is conveyed by a BOOKMARK event when bookmarks are requested. Current upstream documentation describes the feature as Beta since Kubernetes 1.34 and enabled by default there, but distributions can differ. Confirm feature state and flags for the exact API server release before relying on it.

A request can be shaped like this:

GET /api/v1/namespaces/production/pods?watch=true&sendInitialEvents=true&allowWatchBookmarks=true&resourceVersion=&resourceVersionMatch=NotOlderThan

Treat the synthetic initial events as cache population, not as proof that each object was just created. A consumer that triggers a “new Pod” side effect for every ADDED event could duplicate work during every initial synchronization. Track the stream’s initialization boundary and make any downstream action idempotent. Streaming does not remove the need for a bounded cache, backpressure handling, or relisting when the watch cannot resume.

Scope, selectors, and permissions are part of the cache key

A local cache is only meaningful for the API collection it represents. Namespace scope, label selectors, field selectors, resource type, and relevant authorization identity affect what the server returns. If any of these change, the old cache may contain objects no longer in scope or omit objects newly in scope. Rebuild the cache rather than reusing its resource version under a different query.

Grant the controller the narrowest list and watch permissions required for its resource types and namespaces. watch can reveal object data over time just as list can reveal it at one moment. Avoid broad cluster-wide watch permissions when namespace-scoped access is sufficient. Also avoid exposing raw watch streams or cache contents to untrusted users; objects can contain sensitive configuration and status even when their names appear harmless.

Selectors reduce response size and cache churn, but they should reflect durable ownership or workload boundaries. If clients implement selector changes dynamically, test delete/add behavior at the boundary: an object that stops matching may appear as a deletion from the watched collection even though the object itself still exists in the API.

Diagnose a stale or expensive watcher

Measure the cache and its API traffic as one system. Useful signals include list latency and bytes, watch reconnect rate, 410 Gone count, event processing lag, queue depth, cache synchronization duration, client memory, API server request rate, and the rate of full relists. A low reconnect count does not prove freshness if the consumer is blocked processing events; expose a last-applied checkpoint or last-successful synchronization time.

During an incident, compare the local cache with a direct API read under the same namespace and selectors. Check whether the client is connected, whether it has observed a bookmark or other progress recently (without assuming the server must send bookmarks), and whether the event handler has fallen behind. Verify that a 410 path actually rebuilds state rather than merely reopening a stream. For high list cost, narrow the resource scope, paginate, avoid redundant independent watchers, and evaluate supported streaming initial events. Do not increase API server retention expectations or disable fairness controls to hide a client that cannot keep up.

Production acceptance checklist

  • Initial collection state comes from list metadata or a supported streaming-initial-state boundary, not from one object’s resource version.
  • The client applies events idempotently and treats the cache as a recoverable projection rather than an exactly-once event log.
  • A watch resumes from the last safely applied checkpoint; bookmarks are treated as optional progress events.
  • 410 Gone invalidates the cache checkpoint and initiates a complete, correctly scoped relist before the cache is trusted again.
  • Pagination reuses the opaque continuation token and identical query parameters; expired pages restart consistently when atomicity matters.
  • Resource version semantics and streaming-initial-state support are tested on every supported Kubernetes release and distribution.
  • Cache freshness, lag, relists, watch failures, API load, and initialization time are observable and have operational thresholds.
  • RBAC grants only necessary list/watch permissions, and cached objects are protected like API responses.

Correct LIST/WATCH behavior is a foundational reliability property. It prevents missed state after reconnects, protects a controller from treating a partial cache as truth, and reduces unnecessary load on the control plane. Prefer a maintained Kubernetes client library’s reflector or informer implementation unless a custom protocol is required; if building one, test expiration, pagination, restarts, selector changes, slow consumers, and duplicate delivery explicitly.

Related:

Sources:

Comments