Skip to content
SRE & DevOpsDeep Dive Published Updated 6 min readViews unavailable

Kubernetes StorageClasses: Provisioning, Topology, and Reclaim Policy

Design Kubernetes StorageClasses with explicit provisioners, binding topology, default selection, expansion, mount options, and data-retention policy.

A Kubernetes StorageClass describes how a storage provisioner should create persistent volumes for claims that request that class. It is not itself a volume, a backup policy, or a guarantee that a particular workload can mount storage on every node. The class ties together a provisioner, provider-specific parameters, binding timing, topology, mount options, expansion support, and reclaim behavior. Each of those fields can change scheduling or data-retention outcomes.

The safest design is explicit: name the class, document which CSI driver owns it, decide whether provisioning should wait for Pod placement, and state what should happen to the underlying storage after the claim is released. Avoid treating a cluster’s default StorageClass as an invisible global constant. Defaults can change, multiple default classes can exist during migrations, and an omitted class is different from an explicitly named one.

Define a class around an operational contract

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: production-csi-retain
provisioner: csi.example.invalid
reclaimPolicy: Retain
allowVolumeExpansion: true
volumeBindingMode: WaitForFirstConsumer
mountOptions:
  - noatime
parameters:
  performanceTier: balanced

The driver name, parameter keys, and mount options in this sample are placeholders, not portable values. CSI parameters are driver-specific. Kubernetes does not validate every mount option against the target filesystem, so a bad option can cause the volume mount to fail later. Verify the driver’s documentation and test provisioning, mounting, resizing, detach, and reclaim behavior on the actual backend.

The class is cluster-scoped and its name becomes a user-facing request. Treat class names as API contracts for application teams. If balanced or production-csi-retain changes meaning, existing PVCs and newly provisioned PVs can have different characteristics even though their manifests appear unchanged. Prefer versioned class names for materially different provisioning behavior, and communicate deprecation before removing a class in use.

Choose binding mode with scheduler topology in mind

Immediate binding, the default when no mode is set, provisions or binds a volume as soon as the PVC is created. That works well when storage is accessible throughout the cluster. With topology-constrained storage, immediate provisioning can select a zone before the scheduler knows where the consuming Pod can run. The Pod may then be unschedulable because its volume is only available somewhere incompatible with its other constraints.

WaitForFirstConsumer delays binding and dynamic provisioning until a Pod uses the claim. The scheduler can consider node selectors, affinity, anti-affinity, taints, tolerations, and resource requirements together with volume topology. This is often the safer choice for zonal CSI storage when the driver supports it. It is not magic: a node that cannot satisfy all constraints still leaves the Pod pending, and the driver must implement the needed topology behavior.

Do not set spec.nodeName on a Pod that relies on WaitForFirstConsumer; doing so bypasses scheduler placement and can leave the PVC pending. Express node preferences through scheduler-visible constraints such as nodeSelector or affinity, then inspect scheduler events and PVC conditions to understand why binding has not completed.

allowedTopologies can restrict where a class may provision, but overly narrow topology lists reduce placement flexibility and can make recovery harder after a zone loss. Prefer the scheduler-aware binding mode when it represents the intended contract, and use explicit topology constraints only when the platform requires them. Test failover into every supported zone and verify that a restored claim can be placed where its data is accessible.

Make reclaim behavior a data decision

For dynamically provisioned PVs, a StorageClass can set reclaimPolicy to Delete or Retain. If omitted, Kubernetes defaults the class policy to Delete. Delete asks the provisioner to remove the backing storage when the PV is released, subject to driver behavior. Retain leaves the PV in a released state for manual reclamation; an operator must protect, inspect, and eventually clean up the storage.

Neither choice is universally safer. Delete reduces abandoned cloud volumes but can destroy the only copy of data after a claim is removed. Retain protects against some accidental deletions but accumulates storage and needs a documented recovery workflow. The PVC, PV, backend volume, snapshot, and database backup can each have separate lifecycles. A retained PV is not automatically a backup, and a deleted PVC does not imply that the backend data was removed if the class or provisioner behaves differently.

Before applying a reclaim policy, test PVC deletion, namespace deletion, PV release, driver cleanup, and recovery from a retained volume. Capture provider-side identifiers and tags in an inventory. For production data, combine access controls, deletion protection where available, independent backups, and restore tests; a StorageClass alone cannot provide a complete retention guarantee.

Decide defaults and expansion intentionally

A default StorageClass is used for PVCs that omit storageClassName or, in some migration situations, have an empty field. If more than one class is marked default, Kubernetes documentation describes how the most recently created default is selected for new claims without an explicit class. That behavior can turn a routine class rollout into a storage migration with different cost, performance, topology, or deletion semantics.

For durable workloads, specify storageClassName in the PVC when the dependency is intentional. To request no dynamic class, use the documented empty-string behavior rather than leaving the field ambiguous. When changing the default, audit existing manifests, admission defaults, Helm charts, operators, and PVC templates first. Existing bound volumes do not automatically change their provisioner just because a different default class appears.

allowVolumeExpansion: true permits growth for compatible volume types and drivers; it does not make shrinking possible. The claim’s requested capacity can be increased, but the CSI driver, filesystem, and application must support the complete resize path. Monitor PVC conditions and events until filesystem expansion completes. Do not treat an updated requested size as proof that the mounted filesystem already exposes the new capacity.

Diagnose pending claims and failed mounts

When a PVC remains Pending, inspect the claim events, StorageClass, provisioner, binding mode, and the consuming Pod’s scheduling events. With WaitForFirstConsumer, a Pending PVC can be expected until a Pod is scheduled. With Immediate, a missing CSI controller, invalid parameter, exhausted backend quota, or inaccessible topology can prevent provisioning. Read the storage controller and driver logs before deleting objects or repeatedly recreating the claim.

If a Pod has a bound claim but cannot mount it, inspect the CSI node plugin, attach limits, filesystem format, node topology, mount options, and backend state. Storage attachment and mount are separate lifecycle stages; a volume can be provisioned and bound while the Pod still cannot use it. Avoid editing PV objects manually to force a desired status unless the driver’s recovery procedure requires that exact operation.

kubectl get storageclass, kubectl describe pvc, kubectl describe pv, and Pod events provide useful control-plane state. Correlate those with CSI controller and node-plugin logs and provider-side events. Preserve PVC and PV UIDs along with volume handles so a recreated claim with the same name is not mistaken for the original data object.

Test the class before workload rollout

Create a disposable PVC and Pod that exercise provisioning, scheduling, mount, read/write, expansion if enabled, detach, and the selected reclaim policy. Repeat across supported zones and after driver upgrades. Validate that a default change does not alter critical claims, that a retained volume can be safely rebound according to policy, and that a deleted volume is truly cleaned up in the backend.

StorageClass changes affect data placement and lifecycle, not just infrastructure convenience. Keep the class contract explicit, align binding with topology, state the reclaim policy, and test the driver’s real behavior before teams depend on it.

Related:

Sources:

Comments