Vault & secrets interview questions
Prep Vault interviews from KV basics to production ops: auth, policies, dynamic secrets, Kubernetes, leases, seal/unseal, and the secret-zero problem.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Fundamentals
The short version: centralized secrets with identity-based access, leasing, rotation, and audit — so credentials aren't scattered in config files, env vars, and CI, and every access is authenticated and logged.
vault kv put secret/app db_pass=s3cr3t vault kv get -field=db_pass secret/appLink to this question
Engines and auth methods are mounted at API paths — secret/, database/, auth/kubernetes/. Clients read and write those paths; policies grant capabilities per path. The API and the ACL model share the same namespace.
vault secrets list vault kv get secret/appLink to this question
Both store static key/value data; v2 adds versioning, soft delete, undelete, and check-and-set. I'd prefer v2 for new mounts unless you specifically need the simpler v1 semantics.
vault kv put secret/app pass=v2 vault kv get -version=1 secret/app vault kv rollback -version=1 secret/appLink to this question
An auth method — Kubernetes, AppRole, OIDC, cloud IAM — verifies an identity and issues a Vault token bound to policies and a TTL. Clients present that token on later requests, not a shared password.
vault login -method=userpass username=alice # token inherits policies + TTL from the auth roleLink to this question
Vault starts sealed: storage encryption keys aren't in memory. Unsealing recovers the root key (older docs call it the master key): Shamir shares rebuild the unseal key that decrypts it, or auto-unseal has a KMS/HSM decrypt it, so Vault can decrypt storage and serve requests.
Sealed Vault refuses almost all operations; stealing the disk without unseal keys or KMS access is useless. Shamir needs a human quorum after restart; auto-unseal trades that ops burden for trust in the KMS. Lost unseal material is a cluster-loss scenario — protect it and rehearse recovery.
vault operator unseal # enter share 1 vault operator unseal # enter share 2 vault status # Sealed: falseLink to this question
HCL policies grant capabilities — create, read, update, patch, delete, list, sudo — on paths, deny-by-default. Auth roles attach policies to issued tokens. Least-privilege path scoping is the whole ACL model.
path "secret/data/app/*" {
capabilities = ["read"]
}
path "database/creds/app" {
capabilities = ["read"]
}Link to this questioncreate, read, update, patch (partial updates such as vault kv patch), delete, list, and sometimes sudo or deny. list is separate from read — missing list can hide keys even when read on a known path works. deny overrides allow when both match.
path "secret/metadata/app/*" { capabilities = ["list"] }
path "secret/data/app/*" { capabilities = ["read"] }Link to this questionIt's a handle that lets you look up or revoke a token without knowing the token value itself — handy for operators and audit workflows when the raw token has to stay secret.
vault token lookup -accessor <accessor> vault token revoke -accessor <accessor>Link to this question
Vault can map many auth aliases — OIDC user, GitHub, AppRole — to one entity and assign group policies. That keeps ACL consistent when the same human uses different login methods.
Without entities, each auth method issues tokens with only that method's policies — duplicated and inconsistent. Entities unify identity; groups attach shared policies like team-platform. Useful for human SSO plus a break-glass AppRole mapped to the same entity. Machines usually stay on dedicated roles without human entity merging. Review entity merges carefully — a wrong merge elevates access.
vault write identity/lookup/entity name=alice vault read identity/group/name/platform
Interviewer often follows with: How do you stop an OIDC user from inheriting a machine AppRole's policies?
Link to this questionVault from 1.15 onward is under the Business Source License 1.1: production use is allowed, but not offering it to third parties as a hosted or embedded product that competes with the vendor's paid versions. OpenBao is the MPL-2.0 fork run under the Linux Foundation's OpenSSF. For most internal platforms the choice comes down to support, Enterprise-only features, and upgrade path, and legal should read the actual licence text.
Interviewer often follows with: What would you test before pointing existing Vault clients, Terraform providers, and injectors at OpenBao?
Link to this questionDynamic secrets & engines
Vault generates credentials on demand with a lease and TTL — for example a short-lived database user — and revokes them at expiry, so no long-lived shared password sits in config. I'd use them wherever the backend can mint credentials (databases, cloud IAM, SSH) and keep static KV only for what Vault can't create, like third-party API keys, rotated on a schedule and versioned in KV v2.
Each checkout is unique and tied to an identity in the audit log. A leak is time-bounded by TTL; operators can revoke a lease tree instantly during an incident. Needs a secrets engine that can create and delete users or keys in the backend.
vault read database/creds/readonly # username, password, lease_id, lease_durationLink to this question
I'd call it encryption as a service: apps send plaintext and get ciphertext back without holding the key. Transit also signs, verifies, and HMACs; key rotation and rewrap stay inside Vault.
vault write transit/encrypt/orders \ plaintext=$(echo -n 'hi' | base64) vault write -f transit/keys/orders/rotateLink to this question
I'd mount transit, create a named key, grant apps encrypt/decrypt only on that key path, and keep ciphertext in the app database. Keys never leave Vault; I rotate and rewrap on a schedule.
Transit is cryptography as a service, not a KV store — you store ciphertext elsewhere. Policies should split encrypt-only vs decrypt: many writers, few readers. Use convergent encryption only when you understand the dedup/leak tradeoffs. Rotate keys with rewrap so old ciphertext stays readable. For high volume, use batch APIs and be careful about caching tokens. Transit doesn't replace TLS or disk encryption of the DB.
path "transit/encrypt/orders" { capabilities = ["update"] }
path "transit/decrypt/orders" { capabilities = ["update"] }
# prefer separate roles: producers encrypt-only, readers decryptInterviewer often follows with: What happens to old ciphertext after you rotate the transit key?
Link to this questionA lease is Vault's time-bound handle on a dynamic secret or a service token (batch tokens carry no lease). Clients renew it before the TTL runs out, up to max_ttl; when it expires or is revoked, Vault invalidates the credential. Operators can revoke one lease or a whole prefix instantly — a big advantage over static shared secrets.
Renewal takes the lease ID, and Vault Agent usually renews for the app. Once max_ttl is reached, the client has to check out a new secret instead of renewing forever. In an incident, revoking a prefix such as database/creds/readonly revokes every lease issued under that role in one call.
vault lease lookup <lease_id> vault lease renew -increment=30m <lease_id> vault lease revoke -prefix database/creds/readonly # incident: every lease under the role vault token revoke -selfLink to this question
The PKI engine issues short-lived X.509 certs from a configured CA — so services get automated mTLS that rotates frequently instead of long-lived files sitting on disk.
vault write pki/issue/web \ common_name=web.example.com ttl=72hLink to this question
I'd revoke the lease prefix immediately, rotate the database root/config connection if needed, fix the app logging, shorten TTLs, and use transient env or memory-only injection so creds aren't written to disk or stdout.
Revoke first, investigate second. Audit logs show which identity checked out the lease. Root rotation for the database engine updates the admin credential Vault uses without redistributing it to apps. Prevent recurrence: structured logging redaction, agent/CSI file modes with tight permissions, no echo in CI, and TTL short enough that a leaked password dies quickly. Prefer per-role creation statements with least privilege. Post-incident: confirm no persistent clone of the user remains in the DB.
vault lease revoke -prefix database/creds/app vault write database/roles/app \ default_ttl=15m max_ttl=1h \ creation_statements=@readonly.sql
Interviewer often follows with: How do you rotate the database engine's root credentials?
Link to this questionvault secrets enable -path=... mounts an engine; tune sets default and max lease TTLs. Paths become the policy surface — plan mount paths before apps hardcode them.
vault secrets enable -path=secret kv-v2 vault secrets tune -default-lease-ttl=1h -max-lease-ttl=24h database/Link to this question
Kubernetes, CI & apps
The workload presents its ServiceAccount JWT; Vault validates it with TokenReview and maps the SA and namespace to a role and policies. No static Vault token needs to live in the Pod — that's the whole point.
This solves secret zero inside the cluster: the platform already issues the SA token. When Vault runs in the cluster, omit token_reviewer_jwt so it uses its own short-lived pod token; outside the cluster, use the client JWT as the reviewer (clients need system:auth-delegator) or a long-lived reviewer token. Set audience on roles. Bound SA names and namespaces must be tight — a wildcard SA binding is an escape hatch. Prefer projected, audience-bound tokens over long-lived SA secrets. Agent injector or CSI then uses the resulting Vault token to fetch secrets and renew leases.
Pod SA JWT → Vault TokenReview → role/policies → short-lived Vault token.
vault auth enable kubernetes vault write auth/kubernetes/config \ kubernetes_host=https://kubernetes.default.svc vault write auth/kubernetes/role/app \ bound_service_account_names=app \ bound_service_account_namespaces=prod \ policies=app ttl=1h
Interviewer often follows with: What goes wrong if bound_service_account_names is "*"?
Link to this questionInjector adds a Vault Agent sidecar that renders and renews secrets into a shared memory volume; the CSI provider mounts secrets as ephemeral volumes without a sidecar; Vault Secrets Operator (HashiCorp's supported operator) and ESO sync secrets into native Kubernetes Secrets. I'd choose based on whether apps read files or expect Secret objects.
Injector: minimal app change, auto-renew, extra container. CSI: clean mount API, renew depends on driver and version. ESO: easiest for apps already using envFrom or secretKeyRef, but it materializes secrets into etcd — pair with encryption at rest and tight RBAC. VSO takes the same Secret-based approach for Vault shops, handles dynamic and PKI secrets, can restart Deployments or StatefulSets on rotation, and puts the least load on Vault because one operator serves the cluster. Many platforms use ESO for non-sensitive config sync and injector or CSI for high-value dynamic creds.
ESO pulls from Vault (or cloud SM) into a Kubernetes Secret the Pod already consumes.
annotations: vault.hashicorp.com/agent-inject: "true" vault.hashicorp.com/role: "app" vault.hashicorp.com/agent-inject-secret-db: "database/creds/app"
Interviewer often follows with: If ESO syncs into a Secret, what still protects etcd?
Link to this questionTo fetch secrets you need a bootstrap credential — it's an infinite regress. I'd solve it with platform identity — Kubernetes SA, cloud instance identity — that Vault can verify so nothing static is pre-planted.
Planting a long-lived Vault token in a Secret or AMI just moves the problem. Pattern: workload proves who it is via the platform; Vault trusts the platform's attestation; short-lived Vault token follows. For CI, OIDC or carefully delivered wrapped AppRole secret-ids. For VMs, AWS IAM / GCP / Azure auth methods. Document the trust anchors — if the platform IdP is wrong, Vault will mint wrongly.
# no VAULT_TOKEN in the Pod manifest # role bound to SA app in namespace prod
Interviewer often follows with: How do you bootstrap the first Vault admin without secret-zero theater?
Link to this questionFor automated clients outside Kubernetes — CI jobs, VMs, appliances — that authenticate with a RoleID plus a SecretID delivered out-of-band. I'd prefer platform identity when it exists.
vault write auth/approle/role/ci \ token_policies=ci token_ttl=20m vault read auth/approle/role/ci/role-id vault write -f auth/approle/role/ci/secret-idLink to this question
If the agent runs in Kubernetes, I'd use Kubernetes auth bound to that SA. If it's a classic VM agent, AppRole or cloud IAM auth — with wrapped, single-use secret-ids and short token TTLs, never a long-lived root token in Jenkins credentials.
Jenkins Vault Plugin often defaults to a stored token — that's secret zero. Better: AppRole where Jenkins only stores role-id; secret-id is injected per job via a trusted broker, or JWT/OIDC auth if available. Limit policies to the deploy paths needed. Separate roles for PR builders — no prod — vs main. Audit every login. Rotate secret-ids and revoke on agent compromise. Prefer moving the deploy step to a pipeline with OIDC to cloud and Vault JWT auth over a standing agent when you can.
vault write -f -wrap-ttl=60s \ auth/approle/role/ci/secret-id # deliver wrapping token to job; unwrap once
Interviewer often follows with: Why wrap a secret-id instead of writing it to the Jenkins credential store?
Link to this questionYou annotate the Pod with role and secret path; the injector adds an init or sidecar agent that authenticates, writes the secret to a shared volume, and renews leases. The app just reads the file.
vault.hashicorp.com/agent-inject: "true" vault.hashicorp.com/role: "app" vault.hashicorp.com/agent-inject-secret-db: "secret/data/app"Link to this question
I'd check bound_service_account_namespaces and names on the Vault role, the SA token audience and issuer, and whether policies differ — identical role names across mounts still bind explicitly to namespace.
Kubernetes auth roles aren't "same name means access." bound_service_account_namespaces must include the pod namespace; aliases and different auth mounts confuse operators. Also verify TokenReview works for that cluster, projected token expiration, and that the policy path matches secret/data/<ns>/.... Compare vault read auth/kubernetes/role/<name> across envs. Audience mismatches after Kubernetes API changes bite too. Systematic Vault role vs K8s SA binding checks.
vault read auth/kubernetes/role/app # bound_service_account_names # bound_service_account_namespaces kubectl -n other get sa app -o yaml
Interviewer often follows with: How do you safely allow the same app SA name in many namespaces?
Link to this questionUse the JWT auth method with the CI provider's OIDC token. The role checks the issuer, audience, and claims such as repository and branch, then issues a short-lived token with a narrow deploy policy, so nothing long-lived sits in CI settings.
Configure auth/jwt with the provider's OIDC discovery URL and bound_issuer, then one role per trust boundary: role_type set to jwt (the default is oidc), a user_claim, bound_audiences, and bound_claims on repository plus ref or environment so a fork or feature branch can't assume the prod role. Map-valued settings such as bound_claims have to be written as one JSON object, not CLI key=value pairs. Keep token_ttl in minutes and policies scoped to the deploy paths. Prefer explicit claims over sub: providers change subject formats, and GitHub moved new repositories to an ID-based default sub.
vault auth enable jwt
vault write auth/jwt/config oidc_discovery_url="https://token.actions.githubusercontent.com" bound_issuer="https://token.actions.githubusercontent.com"
vault write auth/jwt/role/deploy-prod -<<EOF
{ "role_type": "jwt", "user_claim": "repository",
"bound_audiences": "https://github.com/octo-org",
"bound_claims": { "repository": "octo-org/app", "ref": "refs/heads/main" },
"token_policies": "deploy-prod", "token_ttl": "10m" }
EOFInterviewer often follows with: What stops a pull request from a fork, or a feature branch, from using the prod role?
Link to this questionOperations & HA
Audit devices log every request and response — secrets HMAC'd — to file, syslog, or socket. If Vault can't write audit logs it refuses requests. Auditing is mandatory, not best-effort.
Ship audits to a SIEM; alert on auth failures, policy changes, and unusual path reads. HMAC means raw secrets aren't in the log, but path and identity still enable forensics. Use at least one durable audit device; dual devices avoid a single broken sink blocking the cluster. Protect log integrity — WORM or object lock — so an attacker with Vault admin can't silently erase history on the only copy.
vault audit enable file file_path=/var/log/vault_audit.log vault audit list
Interviewer often follows with: What happens if the audit disk fills up?
Link to this questionFor HA I'd run Integrated Storage — Raft — across an odd number of voting nodes with one leader. Back up Raft snapshots and rehearse restore; Enterprise adds performance/DR replication across sites. And protect the unseal or auto-unseal path through the disaster.
HA keeps serving through node loss; DR is about losing the region. Snapshots must be offline and tested; restoring into a different cluster needs a forced restore plus the unseal keys or KMS access that match the snapshot. Document who holds Shamir shares or how auto-unseal KMS is recovered. Don't treat replication enabled as a backup substitute without restore drills. Monitor raft peer health and leadership flaps.
vault operator raft snapshot save vault.snap vault operator raft snapshot restore vault.snap
Interviewer often follows with: Where do you store Shamir shares relative to the Vault cluster?
Link to this questionConfirm it's sealed vs down, unseal with the quorum or rely on auto-unseal KMS, verify raft peers and the active listener, then check apps retrying auth. If unseal keys aren't available, escalate to DR restore — don't try to "fix" storage blindly.
Distinguish the cases. After a plain process crash, auto-unseal should recover on its own. If the KMS is down or unreachable, Vault stays sealed until it returns: check the KMS IAM permissions, the network path, and whether the key is scheduled for deletion. With Shamir, collect threshold shares through the break-glass procedure, never over chat; a lost quorum is an organizational emergency. The health endpoint and vault status guide automation; after unseal, confirm Raft peers and that audit devices still write. Sealed isn't data loss while storage and unseal material are intact, but availability is zero until unseal. Apps should treat Vault blips as retryable; critical control planes need cached credentials with short remaining TTL only. After recovery, review audit for what failed and whether any break-glass was used, then rehearse unseal, monitor seal status, and keep the list of share holders current.
vault status curl -s https://vault:8200/v1/sys/health vault operator raft list-peers # voters and leader
Interviewer often follows with: How do you test auto-unseal failure without taking prod down?
Link to this questionDon't store unique irreplaceable data only in Vault without backups of recovery material; don't share root tokens; don't disable audit "temporarily"; don't grant sudo on * to app roles.
vault token lookup vault policy read app # expect narrow paths vault audit list # at least one deviceLink to this question
I'd separate two cases. Rotating the KMS key itself (automatic rotation, or a new key behind the configured alias for AWS KMS) needs no Vault downtime as long as old key versions stay enabled. Switching seal type or provider is a seal migration: brief downtime on CE, or online with Seal HA on Enterprise. Either way, never disable or delete a key Vault may still need.
Auto-unseal stores the root key encrypted by the KMS key. For AWS KMS, Vault stores the key information with the encrypted data, so automatic or manual rotation works: new writes use the current key and old keys decrypt older data, which is why they must never be disabled or deleted. Changing seal type is a seal migration: back up, add disabled = true to the old seal block plus the new block, and unseal each node with -migrate; Seal HA on Enterprise does it online and rewraps seal-wrapped values. Test in non-prod. Losing KMS access means a sealed cluster. Seal migration is a practiced DR procedure with provider-specific docs.
vault status # Seal Type: awskms / azurekeyvault / gcpckms # follow provider rewrap/migrate seal docs before disabling old key
Interviewer often follows with: What's your recovery path if the KMS key is scheduled for deletion by mistake?
Link to this questionAccess control & security
Prefer dynamic secrets so rotation is automatic. For static secrets use scheduled rotation and KV versions. Break-glass is a tightly audited, high-privilege path with alerting — used only in emergencies and rotated after.
Break-glass shouldn't be the same as day-two admin SSO. Require dual control where possible, page on use, and time-box tokens. Revoke the initial root token once setup is done; generate one only for break-glass, then revoke it. From Vault 2.0, generate-root and rekey also need a valid Vault token unless enable_unauthenticated_access lists them, so the runbook needs an authenticated operator path as well as key holders. Database root rotation and AWS root IAM rotation keep engine credentials off human laptops.
vault write -force database/rotate-root/my-db
Interviewer often follows with: Should the root token live in the team password manager?
Link to this questionChild tokens are revoked when their parent is; orphan tokens survive parent revocation. Use orphans carefully for long-running automation; revoke-by-prefix during incidents.
Default token treeing lets you revoke a batch of job tokens by revoking the parent. Orphans break that chain, and a login through any non-token auth method (Kubernetes, AppRole, OIDC) already returns an orphan: useful for independence, dangerous if forgotten. Prefer short TTLs and renewable leases over immortal orphans. Periodic tokens renew forever until explicitly revoked — treat them like standing credentials with monitoring.
vault token revoke -mode=path auth/kubernetes/login # or revoke accessor / prefix during incident response
Interviewer often follows with: When is a periodic token justified?
Link to this questionI'd create a dedicated policy with exact path prefixes, attach it only to that auth role, use short TTLs, and deny list where enumeration is a risk. Separate CI roles per environment.
path "secret/data/ci/prod/*" {
capabilities = ["read"]
}
path "auth/token/revoke-self" {
capabilities = ["update"]
}Link to this questionPath prefixes with tight policies work on OSS; Enterprise namespaces give stronger admin isolation per tenant. Either way, separate auth roles, no shared highly privileged tokens, and audit per tenant.
OSS multi-tenancy is mostly policy discipline on shared mounts — a Vault admin still sees everything. Namespaces delegate admin within a boundary and reduce blast radius of policy mistakes. For SaaS platforms, combine namespaces or separate clusters with per-tenant mounts and KPIs on leaked cross-tenant reads. Never give tenants root or unrestricted policy write.
path "secret/data/team-a/*" { capabilities = ["read"] }
path "secret/data/team-b/*" { capabilities = ["deny"] }Interviewer often follows with: When is a dedicated Vault cluster per tenant justified?
Link to this questionIt wraps a secret in a single-use, short-TTL wrapping token so you can hand a courier credential that unwraps once — limiting exposure if the wrapping token leaks after use or expiry.
Common for delivering AppRole secret-ids or one-time bootstrap material. The wrapping token is useless after unwrap or TTL. Combine with tight creation ACL and audit on unwrap. Don't log wrapping tokens. Cubbyhole under the hood — unwrap needs connectivity to Vault.
vault write -f -wrap-ttl=5m \ auth/approle/role/ci/secret-id vault unwrap <wrapping_token>
Interviewer often follows with: What happens if two systems try to unwrap the same wrapping token?
Link to this questionI'd sweep token accessors for their tokens and any other orphan or long-lived ones and revoke them, use audit entries for their entity to find and rotate the secrets they touched, rotate auth methods they controlled, and replace human root-like access with SSO groups, short TTLs, and a break-glass procedure.
Standing privilege is an org failure mode. Hunt: carefully list token accessors, identity entity aliases, and git history of policy changes, and pull the paths the entity read and wrote from audit so you know which secrets to rotate. Revoke accessors, disable orphaned AppRoles, rotate secret-ids, and invalidate OIDC group mappings they influenced. The same cleanup applies to a break-glass root token left alive after an outage: revoke it, rekey if key shares were handled informally to generate it, and generate root again only through a formal ceremony, revoked after use. Preventive: no long-lived human tokens, SSO OIDC with group policies, periodic accessor scans, and dual control for policy write. Late cleanup still matters, because a standing privileged token is a dormant breach. Combine audit forensics with identity hygiene — not only revoke root.
vault list auth/token/accessors vault token lookup -accessor <accessor> # policies, TTL, orphan? vault token revoke -accessor <accessor> vault read identity/entity/id/<id>
Interviewer often follows with: How do you detect newly created orphan tokens in near-real time?
Link to this questionReal-world scenarios
I'd list leases for the database mount, revoke by lease_id or prefix, confirm the DB user is dropped, and fix the client to renew/revoke cleanly or use shorter TTLs plus periodic orphan sweeps.
Leases can outlive processes. List under sys/leases/lookup/database/... and vault lease revoke. Also check the DB for users matching Vault's naming pattern. Root cause is usually no shutdown hook, max_ttl too long, or a lost lease_id. Prefer short TTLs, renewable leases with heartbeats, and an agent or sidecar that manages lifecycle. Orphan leases are standing privilege — treat revocation as incident hygiene.
vault list sys/leases/lookup/database/creds/app/ vault lease revoke database/creds/app/<lease_id> vault lease revoke -prefix database/creds/app/
Interviewer often follows with: How do you revoke all leases for one role without touching other roles?
Link to this questionI'd verify Vault can reach the API server, that the reviewer JWT or SA still works, and that issuer/audience on projected tokens match the auth config — then renew the reviewer credential if it expired.
Vault kubernetes auth calls TokenReview. Failures: expired reviewer token, RBAC removed from reviewer SA, wrong kubernetes_host or CA, API server network policy, or bound_audiences mismatch after a K8s version change. Compare vault read auth/kubernetes/config with the current SA token iss/aud. Prefer short-lived projected tokens and automated rotation of the reviewer JWT. Separate "Vault down" from "auth method misconfigured."
vault read auth/kubernetes/config kubectl auth can-i create tokenreviews --as=system:serviceaccount:vault:reviewer # fix jwt / host / ca; retry: vault write auth/kubernetes/login ...
Interviewer often follows with: Why might only one namespace's pods fail while others succeed?
Link to this questionI'd destroy that secret_id — and rotate the role's secret_id if needed — revoke tokens minted from it via audit, rotate any secrets those tokens could read, and switch delivery to response-wrapping with tight TTLs.
secret_id is a credential. Destroy it, tidy if needed, then hunt accessors in audit for the role. Disable unused AppRoles; prefer pull identity — K8s or OIDC — over long-lived secret_ids. Wrap secret_id with wrap_ttl and single unwrap. Treat a chat paste as public disclosure until proven otherwise.
vault write auth/approle/role/ci/secret-id/destroy \ secret_id=<leaked> vault list auth/approle/role/ci/secret-id # audit: find token accessors → vault token revoke -accessor
Interviewer often follows with: When is rotating role_id also required after a secret_id leak?
Link to this questionI'd create a new key version for encrypt, keep older versions available for decrypt, rewrap ciphertext in a controlled job, then schedule min_decryption_version only after rewrap completes.
Transit rotation is versioned: encrypt uses latest; decrypt uses the version in the ciphertext blob. Rotate with vault write -f transit/keys/foo/rotate, rewrap with /rewrap, track progress, then raise min_decryption_version. Never delete key material prematurely. Apps shouldn't assume key version 1 forever. Rewrap is a data-plane migration with progress metrics, not a one-liner.
vault write -f transit/keys/payments/rotate vault write transit/rewrap/payments ciphertext=$CT vault read transit/keys/payments # check latest_version
Interviewer often follows with: What breaks if you set min_decryption_version before rewrap finishes?
Link to this questionI'd revoke excess leases, lower TTLs and max leases for the role, fix clients that create a user per request, and move to pooled connections with fewer long-lived Vault users.
Each creds/ read can create a DB user or session. Runaways: hot loops, missing lease reuse, or too many replicas times short TTL churn. Mitigate: revoke -prefix, raise DB limits temporarily, set max_ttl and role max leases, use connection pooling, and cache credentials in the app or agent until renewal. Monitor lease count vs DB connections. Dynamic secrets need capacity planning like any pool.
vault lease revoke -prefix database/creds/app/ vault read database/roles/app # lower default_ttl/max_ttl; fix app to reuse until renew
Interviewer often follows with: How do you spot a single noisy client causing lease storms?
Link to this questionI'd check the SecretStore auth — SA / K8s auth role binding — ClusterSecretStore namespace restrictions, ESO controller logs, and whether the ExternalSecret refresh interval or error backoff is hiding Vault 403s.
ESO path: controller authenticates → reads Vault path → writes K8s Secret. Failures often sit in the SecretStore referenced SA, wrong mount path for KV v2 data/, or truncated policies. Align with the kubernetes auth role bindings. Confirm the synced Secret's annotations and timestamps. Stale Secrets are worse than missing ones if apps keep running on old passwords after rotation.
kubectl describe secretstore vault-backend kubectl describe externalsecret app-db kubectl logs -n eso deploy/external-secrets
Interviewer often follows with: How should password rotation be sequenced so pods pick up new values without downtime?
Link to this questionBatch tokens have no accessor and can't be revoked individually, and a batch token from an auth method login is an orphan, so it lives until its TTL. I'd cut what it can reach instead: tighten the policy it carries, which takes effect immediately, revoke the leases it obtained by prefix, rotate the AppRole or OIDC role and downstream secrets, and keep CI batch TTLs short.
Batch tokens are encrypted blobs with no storage entry, no accessor, and no manual revocation; they stop at TTL or when a non-orphan parent is revoked. Design CI with short TTLs, wrapped secret_ids, and OIDC where possible. If compromised: rotate role credentials, invalidate downstream cloud keys the job obtained, and review audit for the token's operations while it was valid. Know batch vs service token tradeoffs before choosing them for CI.
# prefer OIDC/AppRole with 15m TTL vault policy write ci-deploy ci-deploy-reduced.hcl # applies to existing tokens at once vault lease revoke -prefix aws/creds/ci-deploy # leases the job obtained # rotate approle secret_id or OIDC role; batch tokens cannot be revoked
Interviewer often follows with: When would you choose a service token over a batch token for automation?
Link to this questionI'd use destroy for compromised versions, restrict undelete and update in policies, alert on undelete in audit, and treat rotation as write-new-version plus app rollout — not casual undelete.
KV v2 soft delete is reversible; destroy is permanent for a version. Policies should separate delete vs destroy vs undelete. Compromised values need destroy plus rotate downstream. Operationally: document rotation runbooks and block broad update on prod paths. Versioning is a feature and a footgun without policy and audit alerts.
vault kv destroy -versions=3 secret/app vault kv metadata get secret/app # policy: withhold update on secret/undelete/prod/* and secret/destroy/prod/*
Interviewer often follows with: Who should be allowed to undelete in production, if anyone?
Link to this questionRelated
- Cheat sheetHashiCorp Vault cheat sheet
- CourseVault from dev to production
- CourseSecrets management foundations
- CourseAdvanced secrets management
- Field noteVault Kubernetes auth: secrets without static tokens
- Field noteShort-lived TLS certificates from Vault PKI
- Field noteExternal Secrets Operator: sync Vault into Kubernetes
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.