Kubernetes CRD Version Migrations: Conversion, Storage, and Safe Rollouts
Plan safe Kubernetes CRD API-version changes with served and storage flags, conversion webhooks, storedVersions, object rewrites, and rollback gates.
A CustomResourceDefinition (CRD) version migration has two separate jobs: preserve a usable API for clients, and rewrite persisted objects into the new storage representation. Changing a manifest or changing which version is marked storage solves neither automatically. The API server can convert objects when serving a different version, while bytes already stored for existing objects remain at their old version until they are written again or explicitly migrated.
Treat the upgrade as a compatibility protocol across the API server, conversion webhook, controllers, GitOps tooling, admission policy, and every client that reads or writes the resource. A safe rollout keeps old and new clients interoperable for an intentional overlap period, measures migration completion, and only removes an API version after both stored data and callers have moved.
Separate served versions from the storage version
Each entry in spec.versions has a name, a served flag, and a storage flag. served: true makes that version available at its API endpoint; served: false removes that endpoint for clients. Exactly one version must be designated as the storage version. New writes are persisted using the currently selected storage version, regardless of which served version the caller used.
An abbreviated CRD version list might look like this:
spec:
group: example.com
names:
kind: Widget
plural: widgets
singular: widget
scope: Namespaced
versions:
- name: v1beta1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
size:
type: string
- name: v1
served: true
storage: false
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
size:
type: string
This is an excerpt from a CRD, not a complete apply-ready object; real definitions also need metadata, names, and schemas for every served version. Similar schemas do not make versions semantically identical by themselves. If a field is renamed, split, merged, defaulted differently, or given a different meaning, define how that data is represented in both versions before deployment.
The API server converts an object to the version requested by a client when needed. That read response does not rewrite the stored object. With the None conversion strategy, Kubernetes primarily changes the apiVersion and applies schema pruning; it does not understand how to translate a renamed field or a changed meaning. Use a conversion webhook for a real schema transformation.
Design the conversion webhook as an API component
For webhook conversion, the CRD’s spec.conversion points to a webhook endpoint and lists accepted ConversionReview versions. The service must be reachable by the API server, present a certificate trusted by the configured CA bundle, and be running before the CRD begins depending on it. A service reference uses the configured namespace and name; Kubernetes documentation specifies port 443 for the Service, although the backing server can listen on a mapped target port.
conversion:
strategy: Webhook
webhook:
conversionReviewVersions: ["v1"]
clientConfig:
service:
namespace: operators
name: widget-converter
path: /convert
port: 443
caBundle: REPLACE_WITH_BASE64_ENCODED_PEM_CA
The converter receives a batch of objects and a desired API version. It must return the converted objects in the same order and preserve identity. Conversion may mutate labels and annotations, but all other metadata must remain unchanged; attempted changes to identity fields such as name, namespace, or UID are rejected. Test single-object and multi-object requests, both conversion directions, malformed input, webhook timeouts, TLS rotation, and API-server connectivity before switching a storage version.
Conversion needs to be deterministic and semantically reversible wherever the API contract promises that. Suppose v1beta1.spec.size becomes v1.spec.capacity. A v1 object written through v1beta1 must not silently lose a capacity unit or reset the value merely because the older schema cannot express it. Decide how unsupported information is preserved, whether old clients may write after the feature is enabled, and whether validation should reject representations that cannot round-trip. Conversion is not a substitute for versioned API design or upgrade communication.
Treat webhook availability as part of API availability. A request that requires conversion can fail when the webhook is unreachable, has an invalid certificate, returns an unsupported ConversionReview, or responds with malformed data. Avoid making the webhook depend on the custom resource operation whose conversion it must perform. Monitor its availability, latency, error rate, certificate expiry, and API-server-to-Service routing as control-plane dependencies.
Change storage without confusing it with migration
The API server tracks versions that have ever been marked as storage in status.storedVersions. When you designate a new storage version, that field can contain both the old and new version. It is a record of possible persisted versions, not a per-object progress counter. Reading every object through the new endpoint does not prove its on-disk representation changed, because the read path can convert on the fly.
After the CRD declares the new version as storage, new creates and updates are written using it. Existing objects remain persisted at their old version until they are rewritten. The API server does not automatically traverse every object just because the CRD changed. Before disabling or deleting an old served version, migrate all stored objects and verify that the old version can safely be removed.
Kubernetes provides Storage Version Migration for this job, but availability is release- and distribution-dependent. The current upstream documentation lists StorageVersionMigrator as Beta starting in v1.35 and disabled by default, requiring the feature gate on relevant components. Do not assume a hosted cluster exposes the API or migration controller simply because kubectl is new enough. Confirm the control-plane version, feature gate, API discovery, and provider support first.
When available and enabled, a migration request for a namespaced widgets resource can be expressed as:
apiVersion: storagemigration.k8s.io/v1
kind: StorageVersionMigration
metadata:
name: widgets-to-v1
spec:
resource:
group: example.com
resource: widgets
Apply only after the CRD’s new storage version and conversion path are ready. Wait for the migration object’s Succeeded condition, inspect its errors and completion status, and confirm the CRD’s status.storedVersions no longer includes the old version before removing that version. A successful request is evidence from the migration controller, not a reason to skip application-level checks on transformed values.
If Storage Version Migration is unavailable, use a reviewed migration client or controller that lists every object and writes it back through the currently selected storage version while preserving required spec and status. Handle pagination, conflicts, deletes during the scan, throttling, retries, and resources created while the scan is running. Repeat or reconcile until the set is stable. Only after proving every object is rewritten should an administrator remove the old version from the CRD status subresource. Manually editing storedVersions is an assertion about completed work, not the work itself.
Use an expand-migrate-contract rollout
- Inventory consumers and data. List CR instances across all namespaces, controller versions, manifests, GitOps sources, admission rules, backups, and external clients. Record fields whose meaning or default changes, objects per namespace, and any legacy representation that cannot be converted losslessly.
- Deploy the conversion service first. Roll out the webhook, Service, TLS Secret/CA bundle, monitoring, and network rules. Exercise both conversion directions against fixtures that include optional, unknown, boundary, and historical fields. Do not point the CRD at an unready endpoint.
- Expand the CRD. Add the new version with served: true and storage: false, keeping the existing version served and stored. Wait for the CRD to become Established; test GET, LIST, create, update, watch, and server-side apply through each version using representative clients.
- Upgrade writers and readers. Roll controllers and deployment tooling to use the new version while old clients are still supported. Check logs for legacy API usage and validate that fields survive read-modify-write cycles. Keep compatibility for older clients during the agreed overlap, or deliberately reject their writes rather than silently losing information.
- Switch new writes to the target version. After the webhook and consumers are ready, mark the target version storage: true and the old version storage: false. Keep both versions served during migration. Verify the CRD spec and API discovery from the cluster rather than trusting the submitted file alone.
- Migrate persisted objects. Run Storage Version Migration where supported or the controlled rewrite process. Track expected object count, successes, failures, and concurrent writes. Compare sampled objects through both endpoints and validate semantic fields, not just their apiVersion strings.
- Contract only after evidence. Confirm the old version is absent from status.storedVersions, no deployed client requests it, and rollback requirements are met. Then set served: false and eventually remove it from the CRD. Watch audit logs, API errors, conversion metrics, and controller events through the full deprecation window.
Useful inventory commands include:
CRD=widgets.example.com
kubectl get crd "$CRD" -o jsonpath='{range .spec.versions[*]}{.name}{" served="}{.served}{" storage="}{.storage}{"\n"}{end}'
kubectl get crd "$CRD" -o jsonpath='{.status.storedVersions}{"\n"}'
kubectl get widgets.example.com --all-namespaces --chunk-size=200 -o name
kubectl get --raw /apis/example.com/v1 | jq '.resources[].name'
The cluster-wide resource count is an inventory aid, not proof of storage version. CRD status does not expose a reliable per-object storage-version listing to this command. Use migration status and, where a controlled validation requires it, approved storage tooling rather than assuming kubectl get -o yaml reveals the raw persisted encoding.
Make rollback a data plan, not a Git revert
Keeping the old version served can help recover a controller deployment, but it does not reverse data transformations. If v1 introduces a required field with no valid v1beta1 equivalent, switching the storage flag backward can make old readers lossy or invalid. Preserve old schemas and conversion code through the rollback window, and define whether rollback requires a second data migration. Do not remove old fields or conversion branches until every supported consumer and recovery procedure has moved.
Back up the CRD definition and representative custom resources before migration, including status where controllers own it. Test on a restored or disposable cluster with realistic object counts and API-server rate limits. Record object counts before and after, conversion failures, migration completion, controller reconciliation, and important business invariants. A CRD upgrade is complete only when clients, storage, conversion, and recovery all agree on the same contract.
Related:
- Kubernetes Server-Side Apply: Field Ownership, Conflicts, and managedFields
- Kubernetes Controller Reconciliation: Idempotent Desired-State Loops
Sources: