Skip to content
SRE & DevOpsDeep Dive Published Updated 5 min readViews unavailable

Kubernetes StatefulSets: Stable Identity Without Database Guarantees

Design StatefulSets around ordinal identity, headless-Service DNS, per-Pod claims, ordered scale operations, and safe rolling updates.

A StatefulSet manages Pods requiring stable identity or ordered operations. Each replica gets a predictable ordinal, network name, and optional PersistentVolumeClaim (PVC). These primitives help run databases, but they are not database guarantees: StatefulSets do not replicate records, elect leaders, preserve quorum, repair a broken cluster protocol, or ensure durability across storage failure. The application still owns consistency, availability, and recovery.

Declare the network and storage contract

A StatefulSet requires a governing headless Service for per-Pod DNS; create it separately with a selector matching Pod labels. Each claim template creates one PVC per ordinal; the StorageClass and CSI driver define provisioning and attachment.

apiVersion: v1
kind: Service
metadata:
  name: ledger-peer
  labels:
    app: ledger
spec:
  clusterIP: None
  publishNotReadyAddresses: true
  selector:
    app: ledger
  ports:
    - name: peer
      port: 7000
      targetPort: peer
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: ledger
spec:
  serviceName: ledger-peer
  replicas: 3
  selector:
    matchLabels:
      app: ledger
  template:
    metadata:
      labels:
        app: ledger
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: ledger
          image: registry.example.com/ledger:1.0.0
          ports:
            - name: peer
              containerPort: 7000
          volumeMounts:
            - name: data
              mountPath: /var/lib/ledger
  volumeClaimTemplates:
    - metadata:
        name: data
      spec:
        accessModes: ["ReadWriteOnce"]
        storageClassName: fast-block
        resources:
          requests:
            storage: 50Gi

This is illustrative, not a turnkey database. Replace the image and StorageClass with tested choices, verify the CSI driver’s topology and recovery behavior, and ensure the selector matches Pod labels and serviceName names the governing Service. The peer-discovery Service sets publishNotReadyAddresses: true so a member can discover peers before it passes readiness; Kubernetes normally withholds a Pod DNS record until the Pod is Ready. Keep client traffic behind a separate Service without this setting unless routing to unready Pods is explicitly safe.

ReadWriteOnce means read-write mounting by one node, not one Pod; multiple Pods on that node may still access the volume. ReadWriteOncePod can constrain a CSI volume to one Pod, but it is available only with Kubernetes 1.22 or later (stable since 1.29) and compatible CSI sidecars and driver support. Check the current Kubernetes and vendor compatibility requirements. Neither access mode replaces database fencing, replication, or recovery.

Understand what stays stable

With default ordinals, Pods are named ledger-0, ledger-1, and ledger-2. Their DNS names follow ledger-0.ledger-peer..svc.; the cluster domain is configurable, and negative caching may delay records. Use readiness and membership checks instead of assuming immediate DNS resolution.

The name identifies a slot, not an immutable Pod object. A rescheduled ledger-1 can have a new UID but request the same ordinal-associated PVC. The volume must still be available, attachable, and mountable; topology constraints or a Pending claim can block startup. Reusing a volume does not repair its contents or replace application-consistent backups and restore tests.

Scale and termination are operational changes

OrderedReady (the default) creates Pods from low to high and waits for predecessors to become Ready; scale-down removes higher ordinals first after successors shut down. This helps membership-aware applications, but readiness that requires quorum from Pods not yet created can deadlock bootstrap. Probe for safe participation and test recovery.

Parallel pod management relaxes ordering but retains identity. Use it only if the application tolerates concurrent joins, removals, and readiness transitions; do not use it to bypass a stuck rollout without testing protocol safety.

Deleting the StatefulSet does not promise ordered graceful termination. For planned ordered shutdown, scale to zero, wait for completion, then delete the controller. PVCs from volumeClaimTemplates are retained by default. Where the cluster supports and enables StatefulSetAutoDeletePVC, persistentVolumeClaimRetentionPolicy can opt into deletion on scale-down or StatefulSet deletion. That policy interacts with the PV reclaim policy, so verify both before cleanup; see the Persistent Volume lifecycle guide.

Do not force-delete a Pod on an unreachable node until the old process cannot rejoin. Force deletion frees the API name without kubelet confirmation, so a replacement may start while the original is alive. Duplicate identities or storage access can cause data loss.

Roll out code with a recovery boundary

RollingUpdate replaces Pods from highest ordinal downward and waits for each updated Pod to become Ready before its predecessor; minReadySeconds can add a stability interval. Bad images, configuration, or readiness can stall progress. Before further changes, inspect StatefulSet conditions, Pod events, logs, PVC state, and rollout status.

A rolling-update partition stages a canary: ordinals at or above it receive the new template, while lower ones stay on the prior revision. It does not ensure old and new database protocols are compatible. Use the vendor’s schema-change procedure, and verify replication lag, quorum, backup restore, and rollback boundaries before advancing.

During a controlled exercise, inspect controller, Pod, claim, and rollout state together:

kubectl get statefulset ledger
kubectl get pods -l app=ledger -o wide
kubectl get pvc
kubectl describe pod ledger-1
kubectl rollout status statefulset/ledger --timeout=10m

Run these in the workload namespace. If a replacement is Pending, correlate its events with claim binding, PV topology, CSI attach events, and application logs; a replica count alone cannot establish that the correct member and data returned.

Acceptance-test Pod restart, node loss with volume reattachment, scale-up and scale-down, failed readiness during rollout, and staged rollback. Verify Pod names, DNS, PVC-to-ordinal mapping, application membership, and data integrity. Kubernetes coordinates slots; the application and storage system must prove recovery correctness.

Related:

Sources:

Comments