Kubernetes ACME Certificates with cert-manager: Issuance, Challenges, and Renewal
Trace cert-manager Certificate, Order, and Challenge states, choose HTTP-01 or DNS-01 safely, and diagnose issuance or renewal failures.
cert-manager turns certificate issuance into a reconciled Kubernetes workflow. A Certificate expresses the requested names, issuer, and destination Secret. For an ACME issuer, cert-manager creates the downstream request, an ACME Order, and one or more Challenge resources, then presents proof that each requested identifier is under the caller’s control. When issuance succeeds, the certificate and private key are stored in the Secret named by the Certificate; cert-manager also reconciles renewal before expiry.
The difficult incidents are rarely solved by repeatedly deleting the Certificate. First determine which resource in the chain stopped progressing, then verify the selected solver from the perspective of both cert-manager and the external CA. Keep staging and production ACME accounts distinct, protect DNS credentials and account keys, and treat a Ready condition as one important state signal rather than a substitute for verifying the certificate actually served by the application.
Understand issuer scope and the resource chain
An Issuer is namespaced: a Certificate can refer only to an Issuer in its own namespace. A ClusterIssuer is cluster-scoped and can be referenced from any namespace. That scope changes both who can request certificates and where account credentials are stored. For a ClusterIssuer, the ACME account-key Secret referenced by privateKeySecretRef is stored in cert-manager’s cluster resource namespace, which defaults to the cert-manager namespace unless the controller is configured differently. Do not assume that this account Secret lives beside each workload’s TLS Secret.
For an ACME flow, the common reconciliation chain is:
- A
Certificateis created or needs renewal. - cert-manager creates a
CertificateRequestfor the selected issuer. - The ACME issuer creates an
Orderwith one or more identifier authorizations. - The order controller creates
Challengeresources for the DNS names that need validation. - A configured solver presents each challenge; cert-manager performs a self-check before asking the ACME server to validate it.
- After successful validation and issuance, the resulting certificate is written to the
Certificate’s target Secret.
Order and Challenge are controller-managed lifecycle resources. They are not intended to be edited as a manual retry interface. An Order cannot be changed after creation; fix the Certificate or issuer configuration that produced it, then use the documented recovery process if a new order is required. Deleting a challenge or order before correcting the underlying DNS, ingress, credential, or policy issue often recreates the same failure and obscures the original evidence.
Configure a staging issuer and a namespaced certificate
The example below uses the ACME staging endpoint so that routing and reconciliation can be tested without treating a test issuance as a trusted production certificate. Replace the contact address and ingress class with values from the environment. The staging issuer’s account key is separate from the leaf private key generated for the TLS Secret.
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: acme-staging
spec:
acme:
email: [email protected]
server: https://acme-staging-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: acme-staging-account
solvers:
- http01:
ingress:
ingressClassName: public-nginx
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: portal-example-com
namespace: applications
spec:
secretName: portal-example-com-tls
issuerRef:
name: acme-staging
kind: ClusterIssuer
group: cert-manager.io
dnsNames:
- portal.example.com
- www.portal.example.com
ingressClassName is the recommended HTTP-01 ingress configuration for most supported ingress controllers; it was added to cert-manager in version 1.12. If a controller requires the older class field or an existing named Ingress must be edited, follow that controller’s documented compatibility path instead of assuming all ingress implementations create solver routes the same way. A Certificate created by an Ingress annotation (the ingress-shim path) still resolves to the same certificate and ACME resource lifecycle, so troubleshoot the generated Certificate rather than stopping at the Ingress object.
After the staging flow is validated, create a separate production issuer with a distinct account Secret and the production ACME directory URL. Do not change an existing production issuer to staging just to test a route: that changes the issuing authority and can replace a valid production certificate with an untrusted staging chain. Keep issuer names explicit in manifests and promotion workflows so the selected environment is visible in review.
Choose HTTP-01 or DNS-01 from the validation path
HTTP-01 proves control by making a challenge response available at a well-known HTTP path for the requested hostname. The CA reaches the public name, while cert-manager creates temporary solver resources and routes the request to the acmesolver Pod. Let’s Encrypt performs HTTP-01 validation on port 80. Check the complete external path: public DNS, load balancer or gateway, ingress class, route precedence, network policy, and the solver Service and Pod. A route that works only from inside the cluster does not prove that the public CA can retrieve the response.
An ingress controller may route multiple hostnames or wildcard rules through the same address, and a CDN, redirect, WAF, or default backend can change which handler receives /.well-known/acme-challenge/. During an incident, inspect the generated solver Ingress and test the public URL from outside the cluster. Confirm that the intended ingress controller has claimed the solver resource and that its data plane is serving the generated response. Avoid broad rewrites or redirects as a first fix; they can mask the actual route-selection problem.
DNS-01 proves control by publishing a TXT value at _acme-challenge.<name>. It works for wildcard certificate names and for services that do not accept public HTTP traffic, but it requires an automated DNS provider integration or another controlled way to update the record. Give cert-manager only the DNS permissions required for the challenge zone. Where the provider and design allow it, delegate _acme-challenge to a less-privileged zone rather than granting a cluster controller broad access to an entire production DNS account.
cert-manager’s DNS self-check normally uses recursive nameservers from the controller’s /etc/resolv.conf to discover authoritative servers, then queries those authoritative servers. Split-horizon DNS, stale recursive caches, unreachable authoritative servers, and recently changed delegation can make the controller observe a different answer from the external CA. Compare the TXT answer at the authoritative servers and through the resolver path cert-manager uses before changing solver credentials. The dns01-recursive-nameservers and dns01-recursive-nameservers-only controller options alter the self-check path; recursive-only checking can take longer because resolver caches may delay visibility.
By default, cert-manager does not follow CNAME records under _acme-challenge. Delegation through a CNAME requires the configured cnameStrategy: Follow behavior and a solver with permission to update the destination zone. Review the exact record chain before enabling it: a synthesized wildcard CNAME can direct the solver toward an unintended zone. DNS-01 credentials belong in a Kubernetes Secret with narrowly scoped access, and should not be copied into application namespaces merely to make a challenge work.
Diagnose from the first stalled resource
Start with the requested Certificate and its namespace:
kubectl get certificate -n applications
kubectl describe certificate portal-example-com -n applications
kubectl get certificaterequest -n applications
kubectl get orders,challenges -n applications
kubectl get events -n applications
Read the latest condition and event on the earliest resource that is not progressing. If there is no Certificate, inspect the manifest or Ingress annotation that should create it. If the Certificate cannot create a CertificateRequest, inspect issuerRef, namespace, API permissions, and issuer readiness. If the request exists but no Order appears, inspect the issuer condition and the ACME account registration or network error. If an Order has a pending Challenge, inspect its solver selection, presentation status, self-check result, and final ACME authorization error.
For HTTP-01, confirm the solver Ingress class, host, path, Service endpoints, Pod readiness, public address, port 80 reachability, and network-policy path. For DNS-01, inspect the provider error, credential Secret reference, zone selection, delegation chain, and authoritative TXT answer. A challenge’s presented status means the solver presentation step was attempted; it does not prove the external CA can retrieve or resolve the challenge. cert-manager performs a self-check and retries it at a fixed interval while the check fails, so a pending challenge can be evidence of delayed propagation or a reachability problem rather than a need to force another order immediately.
After the self-check succeeds, the external ACME server still performs its own validation. If the challenge moves to an invalid state, use the recorded ACME error reason and validate from the public network path, not only from the controller Pod. Avoid repeated production requests while diagnosing; use staging to test DNS, redirects, gateway attachment, and solver permissions. A split-horizon or NAT-loopback deployment may require the advanced waitInsteadOfSelfCheck solver option, available in cert-manager 1.21 and later, but only after confirming that the external validator can reach the challenge. That option skips cert-manager’s own self-check for a configured wait; it does not repair a bad public path and can simply defer the failed validation.
Verify renewal and the certificate clients receive
cert-manager renews the certificate before expiry according to the Certificate’s effective renewal settings. Monitor Certificate readiness and expiry well before the service’s minimum acceptable renewal window. Also alert on repeated failed issuance events, stale CertificateRequest or Order progress, and a TLS Secret that is absent or not updated as expected. A controller can successfully write a Secret while an ingress, gateway, or application continues serving an older certificate because its reload or reference path is not functioning.
Verify both the Kubernetes state and the client-facing endpoint. Check the Secret’s certificate metadata without printing its private key, inspect the served certificate chain and expiration from an external probe, and confirm that the hostname matches the requested DNS names. Test a renewal in staging with the same solver and routing shape, but do not assume that a successful staging authorization proves production account readiness, production DNS permissions, or production CA reachability.
cert-manager exposes Prometheus metrics from its controller, webhook, and cainjector components; use the metric names and labels documented for the exact cert-manager release in use. Pair relevant certificate-state metrics with Kubernetes events and an external TLS probe. This catches failures at separate boundaries: issuance, Secret update, controller reload, and edge delivery. Retain an accountable owner for DNS credentials, ACME account keys, and renewal alerts so certificate automation does not silently become an unowned cluster dependency.
Renewal timing is configurable and is based on the lifetime of the certificate actually issued, not only the requested duration. By default, cert-manager calculates renewal at two-thirds of that lifetime. spec.renewBeforePercentage can express a relative margin and helps avoid a renewal loop when an issuer returns a shorter lifetime than requested. In cert-manager 1.21 and later, spec.renewal can also constrain attempts to configured renewal windows or explicitly disable automatic renewal for a certificate that is intentionally managed out of band. Review the fail-safe behavior carefully: if no configured window remains before expiry, cert-manager uses the calculated desired renewal time rather than allowing the certificate to expire. These fields do not remove the need to monitor the issued certificate and the endpoint serving it.
cert-manager 1.21 also introduced experimental ACME Renewal Information support behind the ACMEUseARI feature gate. When enabled and supported by the ACME server, this can provide a server-recommended renewal window, for example during a CA-wide revocation or key rollover. Treat it as an opt-in input to renewal policy, not as proof that issuance or deployment succeeded; confirm feature-gate state, issuer support, and served-certificate replacement independently.
Related:
- Kubernetes Pod DNS: Search Domains, Policies, and Resolution Tests
- Kubernetes Gateway API: Separating Infrastructure, Routing, and Application Ownership
Sources:
- cert-manager documentation: ACME Issuers
- cert-manager documentation: ACME Orders and Challenges
- cert-manager documentation: HTTP-01 solver
- cert-manager documentation: DNS-01 solver
- cert-manager documentation: Certificate resources
- cert-manager documentation: Issuer scope
- cert-manager documentation: Troubleshooting
- cert-manager documentation: Prometheus metrics
- cert-manager 1.21 release notes: renewal windows, ACME Renewal Information, and solver self-check option
- Let’s Encrypt documentation: Challenge types