Kubernetes CSI Volume Attachment: Diagnose Stalls, Multi-Attach Errors, and Safe Detach
Trace CSI volume attachment from scheduler intent through VolumeAttachment and driver operations, then recover safely without risking mounted data.
A PersistentVolumeClaim can be bound and a Pod can be scheduled while the storage still is not usable on the selected node. For attachable CSI storage, Kubernetes must coordinate control-plane intent, a VolumeAttachment object, the CSI controller service, and a node-side mount operation. A Pod waiting with a volume-related event is therefore not automatically a scheduler problem or a bad PVC. Diagnose which stage has not converged before trying to move the workload or editing the claim.
This article focuses on the attach/detach control path for CSI volumes. It does not replace driver documentation for cloud API errors, filesystem repair, multipath, or data recovery. Storage operations can have application-level consequences; avoid forcing detaches or deleting controller state until the previous node’s ability to write is understood.
Trace the control path
The scheduler selects a node subject to the Pod’s constraints, volume topology, and binding state. The attach/detach controller determines whether the CSI driver requires an attach operation. When it does, the controller creates a cluster-scoped VolumeAttachment that records the intended PV, node, and attacher. The external-attacher sidecar watches those objects, calls the driver’s CSI controller operation, and reports attachment status and any attach error. After attach completes, the kubelet and node plugin perform the node-side publish and mount work needed by the Pod.
The high-level sequence is:
- PVC is bound to a PV, and the Pod is assigned to a node.
- Kubernetes determines whether the driver requires controller attach.
- A
VolumeAttachmentrepresents the desired attach or detach operation. - The external-attacher coordinates with the CSI controller and updates status.
- The kubelet calls the node-side CSI path and mounts/publishes the volume for the Pod.
Not every CSI driver uses the same attach behavior. The CSIDriver object has an attachRequired setting; a driver that does not require controller attach skips that phase. A missing VolumeAttachment is consequently not proof of failure until you check the driver’s declared behavior and the volume type. A successful VolumeAttachment also does not prove the node-side mount succeeded.
Gather evidence before intervening
Start with read-only inspection. Identify the exact Pod, claim, PV, node, driver, and recent events:
kubectl get pod catalog-0 -n production -o wide
kubectl describe pod catalog-0 -n production
kubectl get pvc catalog-data -n production -o wide
kubectl get pv
kubectl get csidriver
kubectl get volumeattachments.storage.k8s.io -o wide
kubectl get events -A --sort-by=.lastTimestamp
VolumeAttachment is non-namespaced. Match its PV, node, and attacher fields to the workload rather than filtering by the Pod namespace. For the specific attachment, inspect its spec and status:
kubectl describe volumeattachment csi-attachment-example
kubectl get volumeattachment csi-attachment-example -o yaml
kubectl get csinode worker-17 -o yaml
Record spec.attacher, spec.nodeName, spec.source.persistentVolumeName, status.attached, status.attachError, and status.detachError. The API defines attachment status as information populated by the external-attacher completing the operation. The object captures desired intent and the report of a controller-side operation; it is not a direct view into every cloud-provider attachment or filesystem mount.
Use the provider’s supported logs and metrics for the external-attacher, CSI controller, node plugin, and storage backend. Compare timestamps: an old attach error may remain visible after a newer retry or after the Pod moved. Check whether the selected node has the driver registered in CSINode, whether the CSI node plugin is ready there, and whether the driver reports a capacity, topology, or backend issue. Avoid dumping credentials, cloud metadata tokens, or raw storage secrets into incident tickets.
Understand what common symptoms do and do not prove
| Symptom | What it indicates | Next evidence to collect |
|---|---|---|
PVC is Pending |
Binding or provisioning has not completed | PVC events, StorageClass, provisioner logs, topology and capacity |
PVC is Bound, Pod remains Pending |
Binding completed but scheduling or later volume handling may still block progress | Pod scheduling events, node constraints, PV node affinity, topology |
| Pod is scheduled but container is waiting | The node may be waiting for attach, mount, image pull, or another setup step | Pod events, VolumeAttachment, kubelet and CSI node logs |
VolumeAttachment.status.attached is false |
Controller-side attach is not reported complete | attachError, external-attacher logs, CSI controller/backend status |
| Attachment reports attached but mount fails | Controller attach succeeded but node publication, device discovery, or filesystem mount may not have | kubelet and node-plugin logs, CSINode, node health, driver-specific diagnostics |
| Multi-attach or already-attached error | The backend or driver sees a conflicting attachment or an old attachment that has not been safely detached | Current and prior node identity, old node reachability, backend attachment inventory, driver events |
These are branching clues, not automatic diagnoses. A FailedAttachVolume event may refer to a provider quota, invalid identity, unsupported zone, stale attachment, or CSI controller issue. A FailedMount can occur after attachment and may have different causes. Read the complete error and correlate it with the controller and node-side timelines.
Interpret access modes correctly
ReadWriteOnce means a volume can be mounted read-write by one node. It can still be used by multiple Pods on that same node, so it is not a per-Pod single-writer lock. ReadWriteMany and ReadOnlyMany describe supported access patterns, but Kubernetes documents that access modes are used for claim matching and do not generally enforce write protection after a volume has been mounted. Storage backend behavior and application-level coordination still matter.
ReadWriteOncePod is the access mode intended to constrain a supported CSI volume to one Pod cluster-wide. It is stable since Kubernetes v1.29 and requires CSI support and compatible sidecars. Verify the CSI driver, Kubernetes version, and sidecar requirements before adopting or migrating to it. Changing access modes on live data is not a safe generic fix for a multi-attach event.
Access mode is also distinct from filesystem mount mode and from a storage backend’s own attachment limits. A backend may permit only one writer attachment even where multiple Pods on the same node can share the mounted filesystem. Conversely, a filesystem that permits multiple nodes to mount does not guarantee that a particular CSI driver or StorageClass supports that topology. Verify the complete path: PVC/PV mode, driver, backend capability, and application consistency model.
Handle rescheduling and stale attachment carefully
When a node is unhealthy, a replacement Pod can be scheduled elsewhere before the old node has conclusively stopped writing. The detach path may wait for the old node, the CSI controller, or the storage service to report a safe transition. This delay can look like a stuck Pod, but it may be protecting data integrity. Force deletion of a Pod only removes the API object; it does not prove the old kernel or application process can no longer write to the disk.
Before attempting recovery, establish whether the old node is powered off, isolated, or still capable of reaching storage. Check the cloud or storage backend’s attachment inventory and the CSI driver’s documented recovery procedure. If the node is inaccessible but may still run, coordinate a fencing or power-off action through the infrastructure owner before allowing another writer. Do not delete a VolumeAttachment simply to make Kubernetes retry: doing so can hide the control-plane record while the backend attachment remains active.
Use the storage provider’s supported force-detach process only when its prerequisites are met and the previous writer is fenced. A force-detach can make a volume available on another node while buffered writes or filesystem state on the old node are uncertain. For databases and other stateful applications, follow their crash-recovery and replication runbook after the storage layer is safe. Kubernetes attachment success is not equivalent to application-level recovery.
If the old node is healthy and the workload is moving intentionally, wait for the normal termination, unpublish, and detach sequence. Stateful workloads may use ordered shutdown, PodDisruptionBudgets, and StatefulSet identity to reduce disruption, but none of those controls bypasses the storage driver’s safety checks. Observe events and controller logs until the old attachment is gone or marked detached before interpreting the new attach as a stuck operation.
Diagnose topology, quota, and identity failures
CSI storage is often zonal or otherwise topology-constrained. A Pod might be scheduled to a node where an already-bound volume cannot attach. Inspect PV node affinity, StorageClass volumeBindingMode, selected-node annotations on pending claims, and the driver’s documented topology keys. For dynamically provisioned storage, WaitForFirstConsumer lets scheduling constraints participate in volume provisioning; it cannot move an existing zonal disk to a different failure domain by itself.
Provider APIs may reject attach because a node or account reached a disk limit, a volume is in a transition state, the node identity lacks permission, or the requested zone is incompatible. Check quota and provider-side audit records rather than repeatedly deleting and recreating the Pod. Retrying a request that is deterministically unauthorized can create noisy events without changing the outcome.
Also verify the CSI deployment itself: controller and node DaemonSet availability, sidecar version compatibility, leader-election health where applicable, RBAC, cloud identity, network access to the provider API, and driver-specific feature configuration. A cluster upgrade that changes Kubernetes or CSI sidecar versions should be compared against the driver’s supported matrix before assuming a new attach failure is caused by the application.
Recovery checklist
- Capture the Pod, PVC, PV,
VolumeAttachment,CSIDriver,CSINode, and event evidence. - Confirm the Pod’s current node and whether a prior Pod/node can still access the volume.
- Identify whether the driver requires attach and which controller or node phase is failing.
- Correlate Kubernetes timestamps with external-attacher, CSI controller, node-plugin, kubelet, and backend logs.
- Validate topology, access mode, backend attachment limits, node registration, identity, and provider quota.
- Fence an old writer before using a provider-supported force-detach procedure.
- Verify the replacement Pod mounts the intended PV and the application passes its own recovery checks.
- Record the root cause and test the same failure mode in a non-production environment.
The production goal is not to make every attach retry immediate. It is to make the state transition observable, bounded by an explicit recovery procedure, and safe for the data. Preserve the attachment record and backend evidence until you can prove which node is allowed to write.
Related:
- Kubernetes Persistent Volume Lifecycle: Binding, Release, and Reclamation
- Kubernetes CSI Volume Snapshots: Lifecycle, Restore, and Recovery Testing
Sources: