Automating secret rotation without downtime

Rotate database and API credentials on a schedule using overlapping versions so nothing breaks mid-rotation.

May 6, 2025·Updated ·5 min readAdvanced·By SecOpsLog · documentation-verified

Rotation goes wrong at a predictable moment: the old credential is disabled while something still holds it. The something is rarely the application that was in the runbook; it is the nightly report job, the read replica's monitoring user, the third-party integration configured in 2022, or the pod that read its environment at start and has not restarted since. A rotation design is therefore mostly an inventory of readers and an overlap window long enough for every one of them to move, and the tooling (Secrets Manager, Vault, an operator) exists to make the window automatic rather than to make it unnecessary.

Where rotation breaks

FailureWhat happenedThe control
in-flight clients failthe password changed in place; version N died before N+1 was anywheretwo valid versions during the window; disable N only after N+1 is confirmed in use
a forgotten reader breaks at 02:00a cron, a replica user, an integration nobody listedthe inventory, then a period of logging who authenticates with N
pods keep the old valuethe Secret changed in etcd; the process cached the env var at starta reload trigger (a rollout, a file watch, SIGHUP) tied to the change
rotation itself failsthe new password never reached the database, or reached it and never was testeda rotation with explicit test-before-promote steps and an alarm on failure
the window was shorter than the deploythe operator revoked N by the clockmeasure reader lag first; the window is longer than the slowest path, not a round number

Secrets Manager: staging labels are the overlap

Secrets Manager models the overlap with version stages. A rotation Lambda runs four steps: createSecret writes the new value under AWSPENDING; setSecret changes the credential in the database to match it; testSecret uses the pending version to connect; and finishSecret moves AWSCURRENT to the new version and AWSPREVIOUS to the old. A failure at any step leaves AWSCURRENT where it was, which is the property the naive playbook lacks. The alternating users strategy goes further: two database users, one active and one _clone, so the outgoing credential stays valid for readers still holding it while the incoming one is already live. For supported managed services, managed rotation does all of this without a Lambda at all.

rotate.sh
# Lambda rotation on a schedule; the function implements the four steps
aws secretsmanager rotate-secret --secret-id prod/db/app \
--rotation-lambda-arn arn:aws:lambda:eu-west-1:111122223333:function:rotate-postgres \
--rotation-rules 'ScheduleExpression=cron(0 3 ? * SUN *),Duration=2h' # a window, not just an interval
# which version carries which label right now
aws secretsmanager describe-secret --secret-id prod/db/app --query VersionIdsToStages
bash — two versions, both real, for as long as the window needs
aws secretsmanager describe-secret --secret-id prod/db/app --query VersionIdsToStages
{ "6c1f…": ["AWSCURRENT"], "d02a…": ["AWSPREVIOUS"] }
aws secretsmanager get-secret-value --secret-id prod/db/app --version-stage AWSPREVIOUS --query VersionId
"d02a…"
AWSPREVIOUS is still retrievable; with alternating users its database login still works until the next rotation replaces it

Vault: rotate a static user, or stop sharing one

vault-static-role.sh
# the same shared user, rotated by Vault on a schedule; readers fetch the current password
vault write database/static-roles/reporting \
db_name=appdb username=reporting \
rotation_period=24h
# or no shared user at all: each consumer gets its own, with an expiry
vault write database/roles/api \
db_name=appdb default_ttl=1h max_ttl=24h \
creation_statements="CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}'; \
GRANT app_rw TO \"{{name}}\";"

A static role is the honest tool for a user that must keep its name: Vault rotates the password on rotation_period (or a rotation_schedule), and every reader asks Vault for the current value instead of remembering one. The window is then whatever gap exists between Vault's rotation and each reader's next read, which is exactly the number to measure. A dynamic role removes the shared credential and, with it, the calendar: each consumer holds a user that expires by itself, and a leaked one is a one-hour problem. The choice is not ideological; some systems only allow one user, and static roles are how those get rotated without a runbook.

Go deeper in a courseAdvanced secretsDynamic credentials, delivery patterns and rotation that fits a real deploy cadence.View course

The last hop: from a new value to a process that uses it

Every upstream mechanism ends at the same wall. External Secrets Operator rewrites the Kubernetes Secret on its refreshInterval; a process that read DB_PASSWORD from the environment at start does not notice. Mounted files update (subPath mounts excepted), and only an application that re-reads or watches the file benefits. The options are a rolling restart triggered by the Secret change (a reloader controller, or a checksum annotation in the pod template), an in-process reload on a signal or a file watch, or a connection layer that fetches on reconnect and treats an auth failure as the cue to fetch again. The last one holds up best, because it needs no coordination with the rotation at all.

Automate the frequency only after measuring the lag
A fifteen-minute overlap with an hourly ESO refresh and a slow rolling deployment guarantees pods holding the old value when it is disabled. Before scheduling rotation, take one manual rotation and record how long each reader took to move; the window is that number with margin, and the schedule is set from the window, never the other way round.

Rotation is easiest for the credentials nobody stores: CI that assumes a role through OIDC federation has nothing to rotate, and a pod that fetches a leased database user through Kubernetes auth rotates by expiring. Where a static value has to exist, ESO makes the copy in the cluster follow the source; the reload is still yours.

Related posts

Quick reference