Short-lived TLS certificates from Vault PKI

Run an internal CA in Vault and issue certs that live for hours, not years — rotation becomes a non-event.

Nov 4, 2025·Updated ·5 min readAdvanced·By SecOpsLog · documentation-verified

Internal TLS fails in two ways, and both come from certificates that live too long. The three-year certificate expires on a weekend with no record of who issued it; the compromised one keeps working because nobody can revoke what nobody tracks. Vault's PKI secrets engine addresses both by making issuance an authenticated API call with a role-enforced ceiling on lifetime and names, so a service can hold a certificate for a day and ask for another. What it does not do, despite what many guides say, is attach a lease to those certificates or kill them when a timer runs out. Understanding what actually ends a certificate's life is most of operating this well.

Root outside, intermediate inside

pki-int.sh
vault secrets enable -path=pki_int pki
vault secrets tune -max-lease-ttl=43800h pki_int # the mount ceiling: 5 years for the intermediate itself
# CSR from Vault; the root (offline, or a separate tightly-policed mount) signs it
vault write -field=csr pki_int/intermediate/generate/internal \
common_name="Acme Internal Issuing CA" key_type=ec key_bits=384 > int.csr
# ... sign int.csr with the root, producing int.pem ...
vault write pki_int/intermediate/set-signed certificate=@int.pem
vault write pki_int/config/urls \
issuing_certificates="https://vault.acme.internal:8200/v1/pki_int/ca" \
crl_distribution_points="https://vault.acme.internal:8200/v1/pki_int/crl"

The intermediate is what Vault issues from; the root signs intermediates and nothing else, and lives where its private key cannot be read by an application role: an offline CA, an HSM-backed mount, or at minimum a separate Vault mount whose policy is held by two people. A lab can generate a root inside Vault with pki/root/generate/internal, and the only mistake is letting that lab root become the trust anchor the fleet pins. config/urls matters more than it looks: the AIA and CRL URLs are baked into every leaf, so clients that check revocation need them reachable, and changing them later means reissuing.

Whoever can write to the issuer can impersonate every internal service
pki_int/intermediate/set-signed, pki_int/root/* and the role definitions are the keys to the kingdom, and they should sit behind a policy that no application role shares, with the Vault audit log alerting on any write to those paths. Keep issuing roles read-and-issue only, and keep the root where an application token cannot reach it at all.

A role that refuses what it should

pki-role.sh
vault write pki_int/roles/shop-svc \
allowed_domains="shop.svc.cluster.local,shop.acme.internal" \
allow_subdomains=true \
allow_bare_domains=false \
enforce_hostnames=true \
key_type=ec key_bits=256 \
ttl=24h max_ttl=72h \
no_store=false # keep the serial so it can be revoked; set true only for very high volume

allowed_domains with allow_subdomains bounds the names, and enforce_hostnames (on by default) rejects a CN that is not a hostname at all. max_ttl is the ceiling a requester cannot argue with, and a role that permits * with a 90-day maximum has rebuilt the long-lived certificate inside Vault with extra steps. One role per service class keeps a compromised AppRole from minting names it has no business holding: shop-svc for workloads, ingress-edge for public hostnames with a different approver.

bash — issue, and read what came back
vault write -format=json pki_int/issue/shop-svc common_name=checkout.shop.svc.cluster.local ttl=24h > cert.json
jq -r ".data | keys[]" cert.json
ca_chain certificate expiration issuing_ca private_key private_key_type serial_number
jq -r .lease_duration cert.json
0
no lease: generate_lease is false by default. The certificate is valid until its Not After, and Vault will not end it for you

What ends a certificate, and who renews it

Lifetime and renewal ownership

MechanismWhat it doesWho is responsible
Not Afterthe certificate stops validating; nothing in Vault has to happenwhoever asked for it must ask again before then
generate_lease=trueattaches a lease; vault lease revoke adds the serial to the CRL when the lease endsVault; costs storage and slows startup at scale, which is why the default is off
pki_int/revokeadds a serial to the CRL (and OCSP) on demandthe operator, on compromise; needs the serial, which no_store=true throws away
Vault Agent pkiCertre-issues when the rendered certificate reaches lease_renewal_threshold (0.9) of its lifetime, treating Not After as the endthe sidecar; the application must reload the file
cert-manager Vault issuerrequests through the same role and rotates the Secret before expirythe controller; ingress and Gateway consumers reload on Secret change

The consequence of no lease is that renewal is the client's job, and the job has an owner or it has an outage. Vault Agent with the pkiCert template function is the in-pod answer: it fetches on start, re-renders at ninety percent of the certificate's life by default, and can run a command to make the process reload. cert-manager's Vault issuer is the Kubernetes-native answer for ingress and service certificates, using the same role and policy. For anything outside a cluster, a cron that calls pki_int/issue and reloads the service is legitimate, provided the expiry is also on a dashboard. Revocation stays manual and needs the serial number, which is the argument for leaving no_store at its default until volume forces the question.

Short certificates are only as safe as the identity that requests them, so this mount is usually fronted by Kubernetes auth and a pod can call issue without a stored token. Where the consumer is an ingress controller, cert-manager and TLS on ingress is the shape that keeps the private key out of a pipeline log.

Related posts

Quick reference