Automating secret rotation without downtime
Rotate database and API credentials on a schedule using overlapping versions so nothing breaks mid-rotation.
Rotation goes wrong at a predictable moment: the old credential is disabled while something still holds it. The something is rarely the application that was in the runbook; it is the nightly report job, the read replica's monitoring user, the third-party integration configured in 2022, or the pod that read its environment at start and has not restarted since. A rotation design is therefore mostly an inventory of readers and an overlap window long enough for every one of them to move, and the tooling (Secrets Manager, Vault, an operator) exists to make the window automatic rather than to make it unnecessary.
Where rotation breaks
| Failure | What happened | The control |
|---|---|---|
| in-flight clients fail | the password changed in place; version N died before N+1 was anywhere | two valid versions during the window; disable N only after N+1 is confirmed in use |
| a forgotten reader breaks at 02:00 | a cron, a replica user, an integration nobody listed | the inventory, then a period of logging who authenticates with N |
| pods keep the old value | the Secret changed in etcd; the process cached the env var at start | a reload trigger (a rollout, a file watch, SIGHUP) tied to the change |
| rotation itself fails | the new password never reached the database, or reached it and never was tested | a rotation with explicit test-before-promote steps and an alarm on failure |
| the window was shorter than the deploy | the operator revoked N by the clock | measure reader lag first; the window is longer than the slowest path, not a round number |
Secrets Manager: staging labels are the overlap
Secrets Manager models the overlap with version stages. A rotation Lambda runs four steps: createSecret writes the new value under AWSPENDING; setSecret changes the credential in the database to match it; testSecret uses the pending version to connect; and finishSecret moves AWSCURRENT to the new version and AWSPREVIOUS to the old. A failure at any step leaves AWSCURRENT where it was, which is the property the naive playbook lacks. The alternating users strategy goes further: two database users, one active and one _clone, so the outgoing credential stays valid for readers still holding it while the incoming one is already live. For supported managed services, managed rotation does all of this without a Lambda at all.
# Lambda rotation on a schedule; the function implements the four stepsaws secretsmanager rotate-secret --secret-id prod/db/app \--rotation-lambda-arn arn:aws:lambda:eu-west-1:111122223333:function:rotate-postgres \--rotation-rules 'ScheduleExpression=cron(0 3 ? * SUN *),Duration=2h' # a window, not just an interval# which version carries which label right nowaws secretsmanager describe-secret --secret-id prod/db/app --query VersionIdsToStages
aws secretsmanager describe-secret --secret-id prod/db/app --query VersionIdsToStages{ "6c1f…": ["AWSCURRENT"], "d02a…": ["AWSPREVIOUS"] }aws secretsmanager get-secret-value --secret-id prod/db/app --version-stage AWSPREVIOUS --query VersionId"d02a…"AWSPREVIOUS is still retrievable; with alternating users its database login still works until the next rotation replaces itVault: rotate a static user, or stop sharing one
# the same shared user, rotated by Vault on a schedule; readers fetch the current passwordvault write database/static-roles/reporting \db_name=appdb username=reporting \rotation_period=24h# or no shared user at all: each consumer gets its own, with an expiryvault write database/roles/api \db_name=appdb default_ttl=1h max_ttl=24h \creation_statements="CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}'; \GRANT app_rw TO \"{{name}}\";"
A static role is the honest tool for a user that must keep its name: Vault rotates the password on rotation_period (or a rotation_schedule), and every reader asks Vault for the current value instead of remembering one. The window is then whatever gap exists between Vault's rotation and each reader's next read, which is exactly the number to measure. A dynamic role removes the shared credential and, with it, the calendar: each consumer holds a user that expires by itself, and a leaked one is a one-hour problem. The choice is not ideological; some systems only allow one user, and static roles are how those get rotated without a runbook.
The last hop: from a new value to a process that uses it
Every upstream mechanism ends at the same wall. External Secrets Operator rewrites the Kubernetes Secret on its refreshInterval; a process that read DB_PASSWORD from the environment at start does not notice. Mounted files update (subPath mounts excepted), and only an application that re-reads or watches the file benefits. The options are a rolling restart triggered by the Secret change (a reloader controller, or a checksum annotation in the pod template), an in-process reload on a signal or a file watch, or a connection layer that fetches on reconnect and treats an auth failure as the cue to fetch again. The last one holds up best, because it needs no coordination with the rotation at all.
Rotation is easiest for the credentials nobody stores: CI that assumes a role through OIDC federation has nothing to rotate, and a pod that fetches a leased database user through Kubernetes auth rotates by expiring. Where a static value has to exist, ESO makes the copy in the cluster follow the source; the reload is still yours.