Kubernetes Ingress TLS with cert-manager and Let's Encrypt

Issue and auto-renew TLS certificates for Ingress with cert-manager, ACME, and a ClusterIssuer, and debug stuck orders before ACME rate-limits you.

Mar 25, 2025·Updated ·5 min readIntermediate·By SecOpsLog · documentation-verified

cert-manager on a good day is invisible: an Ingress declares a tls block, a Secret with tls.crt and tls.key appears, and the certificate renews itself forever. On a bad day a Certificate sits at READY False for an hour, the browser keeps warning, and every retry quietly spends a Let's Encrypt rate limit. The difference between the two days is knowing which of five objects to read, in which order, because the failing step names itself in exactly one of them.

The object chain, and what a stall in each one means

ObjectCreated byStalls when
Ingressyouthe tls block or issuer annotation is missing, so nothing else is created
Certificatecert-manager (ingress-shim) from the Ingressthe referenced issuer does not exist or is not ready
CertificateRequestcert-manager from the Certificatethe issuer rejects the request (wrong ACME account, private key problems)
Orderthe ACME issuerLet's Encrypt cannot create the order: rate limit, invalid identifier, account issues
Challengethe Ordervalidation fails: the /.well-known/acme-challenge/ URL is unreachable, DNS is wrong, the solver targets the wrong ingress class
bash — read the chain from the top until something says why
kubectl get certificate,certificaterequest,order,challenge -n prod
certificate.cert-manager.io/shop-tls False shop-tls 38m
certificaterequest.cert-manager.io/shop-tls-1 False ... 38m
order.acme.cert-manager.io/shop-tls-1-2019 pending 38m
challenge.acme.cert-manager.io/shop-tls-1-2019-1 pending shop.example.com 38m
kubectl describe challenge shop-tls-1-2019-1 -n prod | grep -A2 Reason
Reason: Waiting for HTTP-01 challenge propagation: wrong status code '404', expected '200'
the ACME server reached the host and got a 404: the solver ingress is not receiving the request

An issuer that points at the right ingress controller

For HTTP-01, cert-manager creates a temporary Ingress that routes /.well-known/acme-challenge/<token> to a solver pod. That Ingress must be picked up by the controller that actually receives internet traffic for the hostname; ingressClassName on the solver is what selects it (the older class field survives for controllers that still use the annotation). A 404 on the challenge URL, as above, almost always means the solver Ingress went to a class no external controller serves. Start against the staging endpoint: its certificates are untrusted by browsers, and its rate limits are separate from production's.

cluster-issuer.yaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-staging # debug here; a letsencrypt-prod twin uses acme-v02
spec:
acme:
email: sre@example.com
server: https://acme-staging-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-staging-account
solvers:
- http01:
ingress:
ingressClassName: nginx # the class of the controller facing the internet
ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: shop
namespace: prod
annotations:
cert-manager.io/cluster-issuer: letsencrypt-staging
spec:
ingressClassName: nginx
tls:
- hosts: [shop.example.com]
secretName: shop-tls # cert-manager writes tls.crt and tls.key here
rules:
- host: shop.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: shop
port:
number: 80

Validate the challenge the way the ACME server does

Let's Encrypt fetches the challenge URL over plain HTTP from several vantage points on the public internet. So the test is the same request from outside the cluster: curl -v http://shop.example.com/.well-known/acme-challenge/<token> while the Challenge is pending. A connection refused means DNS or the load balancer; a 404 means the solver Ingress is not being served by the controller behind that address; a redirect to HTTPS is a controller or WAF rule that must exempt the challenge path; a 200 with the token means the problem is on the ACME side, which the Order's status will say.

bash — the four outcomes of the external check
curl -sv http://shop.example.com/.well-known/acme-challenge/probe 2>&1 | grep -E "^< HTTP|Could not resolve|Connection refused"
curl: (6) Could not resolve host -> DNS: no A/AAAA record, or it points elsewhere
< HTTP/1.1 404 Not Found -> solver Ingress not served by this controller
< HTTP/1.1 308 Permanent Redirect -> HTTPS redirect or WAF catches the challenge path
< HTTP/1.1 200 OK -> the path works; read the Order if still pending

The rate limits that turn a typo into a week

Production Let's Encrypt allows 5 certificates for the exact same set of hostnames per 7 days, refilling one every 34 hours, and 50 certificates per registered domain per 7 days. Repeated failed validations for one hostname are limited as well, and the counter resets only when a validation for that name succeeds. A retry loop against a broken solver does not hit the first limit (no certificate is issued) but it does consume the failure allowance, and a working solver that issues once per debugging attempt hits the duplicate limit by lunch. Both are the reason to debug on staging and switch the issuer annotation only when kubectl get challenge has shown valid there.

Once issued, cert-manager renews at two-thirds of the certificate's lifetime by default, with spec.renewBefore on the Certificate to change that (as a Go duration such as 360h, not 15d). Renewal reuses the same solver, so a change that breaks HTTP-01 shows up weeks later as a renewal failure; the Certificate's Ready condition and the certmanager_certificate_expiration_timestamp_seconds metric are what to alert on.

HTTP-01 cannot issue wildcards or internal names
A wildcard certificate, or a hostname that only resolves inside the VPC, needs the dns-01 solver against your DNS provider (Route 53, Cloudflare and others have built-in solvers) because the ACME server must be able to validate without reaching the service. Hostnames that should never appear in public certificate transparency logs belong on an internal CA, which is the case for Vault PKI rather than for Let’s Encrypt.
Symptoms and the object to read
Symptom
Certificate False, no CertificateRequest
Order pending, Challenge pending
Challenge failed: 404 or connection refused
Order errored: rateLimited
Read
kubectl describe certificate: issuer missing or not Ready
kubectl describe challenge: Reason line
curl the challenge URL from outside
wait it out on prod; keep debugging on staging

TLS at the edge covers the browser-to-cluster hop. For the hops inside the cluster the same tooling issues internal certificates from a private CA (Vault PKI), and for keeping a compromised pod from reaching services it has no business with, NetworkPolicy is the control that TLS does not replace.

Related posts

Quick reference