Kubernetes Ingress TLS with cert-manager and Let's Encrypt
Issue and auto-renew TLS certificates for Ingress with cert-manager, ACME, and a ClusterIssuer, and debug stuck orders before ACME rate-limits you.
cert-manager on a good day is invisible: an Ingress declares a tls block, a Secret with tls.crt and tls.key appears, and the certificate renews itself forever. On a bad day a Certificate sits at READY False for an hour, the browser keeps warning, and every retry quietly spends a Let's Encrypt rate limit. The difference between the two days is knowing which of five objects to read, in which order, because the failing step names itself in exactly one of them.
The object chain, and what a stall in each one means
| Object | Created by | Stalls when |
|---|---|---|
| Ingress | you | the tls block or issuer annotation is missing, so nothing else is created |
| Certificate | cert-manager (ingress-shim) from the Ingress | the referenced issuer does not exist or is not ready |
| CertificateRequest | cert-manager from the Certificate | the issuer rejects the request (wrong ACME account, private key problems) |
| Order | the ACME issuer | Let's Encrypt cannot create the order: rate limit, invalid identifier, account issues |
| Challenge | the Order | validation fails: the /.well-known/acme-challenge/ URL is unreachable, DNS is wrong, the solver targets the wrong ingress class |
kubectl get certificate,certificaterequest,order,challenge -n prodcertificate.cert-manager.io/shop-tls False shop-tls 38mcertificaterequest.cert-manager.io/shop-tls-1 False ... 38morder.acme.cert-manager.io/shop-tls-1-2019 pending 38mchallenge.acme.cert-manager.io/shop-tls-1-2019-1 pending shop.example.com 38mkubectl describe challenge shop-tls-1-2019-1 -n prod | grep -A2 ReasonReason: Waiting for HTTP-01 challenge propagation: wrong status code '404', expected '200'the ACME server reached the host and got a 404: the solver ingress is not receiving the requestAn issuer that points at the right ingress controller
For HTTP-01, cert-manager creates a temporary Ingress that routes /.well-known/acme-challenge/<token> to a solver pod. That Ingress must be picked up by the controller that actually receives internet traffic for the hostname; ingressClassName on the solver is what selects it (the older class field survives for controllers that still use the annotation). A 404 on the challenge URL, as above, almost always means the solver Ingress went to a class no external controller serves. Start against the staging endpoint: its certificates are untrusted by browsers, and its rate limits are separate from production's.
apiVersion: cert-manager.io/v1kind: ClusterIssuermetadata:name: letsencrypt-staging # debug here; a letsencrypt-prod twin uses acme-v02spec:acme:email: sre@example.comserver: https://acme-staging-v02.api.letsencrypt.org/directoryprivateKeySecretRef:name: letsencrypt-staging-accountsolvers:- http01:ingress:ingressClassName: nginx # the class of the controller facing the internet
apiVersion: networking.k8s.io/v1kind: Ingressmetadata:name: shopnamespace: prodannotations:cert-manager.io/cluster-issuer: letsencrypt-stagingspec:ingressClassName: nginxtls:- hosts: [shop.example.com]secretName: shop-tls # cert-manager writes tls.crt and tls.key hererules:- host: shop.example.comhttp:paths:- path: /pathType: Prefixbackend:service:name: shopport:number: 80
Validate the challenge the way the ACME server does
Let's Encrypt fetches the challenge URL over plain HTTP from several vantage points on the public internet. So the test is the same request from outside the cluster: curl -v http://shop.example.com/.well-known/acme-challenge/<token> while the Challenge is pending. A connection refused means DNS or the load balancer; a 404 means the solver Ingress is not being served by the controller behind that address; a redirect to HTTPS is a controller or WAF rule that must exempt the challenge path; a 200 with the token means the problem is on the ACME side, which the Order's status will say.
curl -sv http://shop.example.com/.well-known/acme-challenge/probe 2>&1 | grep -E "^< HTTP|Could not resolve|Connection refused"curl: (6) Could not resolve host -> DNS: no A/AAAA record, or it points elsewhere< HTTP/1.1 404 Not Found -> solver Ingress not served by this controller< HTTP/1.1 308 Permanent Redirect -> HTTPS redirect or WAF catches the challenge path< HTTP/1.1 200 OK -> the path works; read the Order if still pendingThe rate limits that turn a typo into a week
Production Let's Encrypt allows 5 certificates for the exact same set of hostnames per 7 days, refilling one every 34 hours, and 50 certificates per registered domain per 7 days. Repeated failed validations for one hostname are limited as well, and the counter resets only when a validation for that name succeeds. A retry loop against a broken solver does not hit the first limit (no certificate is issued) but it does consume the failure allowance, and a working solver that issues once per debugging attempt hits the duplicate limit by lunch. Both are the reason to debug on staging and switch the issuer annotation only when kubectl get challenge has shown valid there.
Once issued, cert-manager renews at two-thirds of the certificate's lifetime by default, with spec.renewBefore on the Certificate to change that (as a Go duration such as 360h, not 15d). Renewal reuses the same solver, so a change that breaks HTTP-01 shows up weeks later as a renewal failure; the Certificate's Ready condition and the certmanager_certificate_expiration_timestamp_seconds metric are what to alert on.
TLS at the edge covers the browser-to-cluster hop. For the hops inside the cluster the same tooling issues internal certificates from a private CA (Vault PKI), and for keeping a compromised pod from reaching services it has no business with, NetworkPolicy is the control that TLS does not replace.