Terraform remote state: S3 native locking, migration, and recovery

Move Terraform state off your laptop into an encrypted S3 or GCS backend with locking, so two applies never race and corrupt your infrastructure.

May 5, 2026·Updated ·6 min readIntermediate·By SecOpsLog · documentation-verified

terraform.tfstate is the inventory of every resource id a stack owns, and for many providers it also holds the plaintext of whatever those resources were created with: database passwords, API keys, TLS private keys. Left on a laptop it has two problems. It is a secret that is not protected like one, and it is the only copy, so a second engineer running apply from their own copy produces two histories of the same infrastructure. Remote state fixes both: one encrypted object everyone reads, and a lock so only one writer changes it at a time.

On current Terraform the lock is an S3 object, not a DynamoDB table. use_lockfile = true on the s3 backend writes <key>.tflock next to the state for the duration of an operation; HashiCorp marks the older dynamodb_table locking as deprecated, and new backends should not create the table at all. What follows is the backend block, the IAM statement that makes it least-privilege, the one-time migration, and what to do when a lock is left behind.

The backend block

Encode environment and stack in the key so the production network and the staging network can never share a state file. encrypt = true asks S3 to apply SSE-S3 or SSE-KMS to the state and lock objects. Turn on bucket versioning before the first production apply, not after: a bad state write is recoverable from the previous object version only if versioning was already on when the write happened.

backend.tf
terraform {
backend "s3" {
bucket = "acme-tfstate"
key = "prod/network/terraform.tfstate"
region = "eu-west-1"
use_lockfile = true
encrypt = true
}
}

IAM: three objects, three permission sets

The documented minimum is s3:ListBucket on the bucket (scoped with an s3:prefix condition to the state path), s3:GetObject and s3:PutObject on the state object, and s3:GetObject, s3:PutObject and s3:DeleteObject on the .tflock object. Terraform never deletes the state object, so a CI role does not need delete on it; a policy that grants s3:* on the bucket hands the pipeline the ability to destroy the history that versioning was meant to keep. Human break-glass roles can be narrower still, limited to non-production prefixes until an incident widens them deliberately.

state-iam.json
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::acme-tfstate",
"Condition": { "StringLike": { "s3:prefix": "prod/network/*" } }
},
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::acme-tfstate/prod/network/terraform.tfstate"
},
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
"Resource": "arn:aws:s3:::acme-tfstate/prod/network/terraform.tfstate.tflock"
}
]
}

Migrate the local file once

Add the backend block, run terraform init -migrate-state, and answer yes to copying the existing state. Afterwards the local terraform.tfstate is stale by definition; delete it and add .tfstate to .gitignore if it is not there already. Listing object versions on the new key is a cheap way to confirm that versioning is really on before anyone relies on it.

bash — migrate and prove versioning works
terraform init -migrate-state
Do you want to copy existing state to the new backend? yes
Successfully configured the backend "s3"!
aws s3api list-object-versions --bucket acme-tfstate \
--prefix prod/network/terraform.tfstate --query "Versions[].[VersionId,LastModified]"
[["3sL4kqtJlcpXroDTDmJ+rmSpXd3dIbrHY", "2026-09-12T09:14:02+00:00"]]
one version now; every apply adds another, and any of them can be restored

What the lockfile does, and the stale-lock decision

One apply, two outcomes

The lock object exists only while an operation runs. The error path is the one to read carefully: the message names the holder, and that name decides whether waiting or force-unlock is the right move. Simplified: the exact S3 request Terraform uses to create the object is an implementation detail and is not shown.

Terraform S3 backend lockfile: apply creates <key>.tflock, runs plan and apply and deletes the lock; if the object already exists the run stops with a state lock error, and force-unlock is only for a holder that is confirmed dead terraform applyuse_lockfile = truecreate the lock object<key>.tflock, fails if presentcreatedalready existsLock held: run the operation1read terraform.tfstate2plan, then apply3write the new state version4delete <key>.tflockError acquiring the state lockIDlock id, who, when, operationthenwait for the other run and retrynotforce-unlock while it may liveBucket versioningevery write keeps the previousstate version for recoveryterraform force-unlock <ID>only once the holder is known dead:a crashed job, not a slow one

A second plan or apply against a locked state fails immediately with Error acquiring the state lock, and the message includes the lock id, who created it, when, and which operation. In a healthy team that error means someone else is mid-apply: wait, then retry. It means something different when the holder is a CI job that was killed ten minutes ago, because a crashed process never deletes its .tflock. That is the only situation terraform force-unlock <ID> is for. Forcing a lock whose holder is still running recreates the exact concurrent-write problem the lock exists to prevent, so confirm the run is dead (the pipeline UI, the runner's process list) before forcing anything.

bash — a lock left behind by a killed job
terraform plan
Error: Error acquiring the state lock
Lock Info: ID: 7f0c1d2e-… Operation: OperationTypeApply
Who: runner@ci-agent-14 Created: 2026-09-12 08:41:07 UTC
ci-agent-14 job #4471 was cancelled at 08:43; nothing else holds the lock
terraform force-unlock 7f0c1d2e-…
Terraform state has been successfully unlocked

Backends that still carry dynamodb_table can keep it while runners are upgraded: Terraform accepts dynamodb_table and use_lockfile together during the migration window, and the deprecated argument is removed once every caller is on a version that supports the lockfile. HCP Terraform, GCS and Azure Blob have their own lock semantics, so a team that spans backends should write down which mechanism each stack uses rather than assuming the S3 rules apply.

The state bucket is a secrets store
Reads on the state object should be limited to CI roles and named break-glass humans, because the object contains whatever the providers wrote into it. Keep .tfstate out of Git entirely: secret scanners find committed state files, and so do the people the scanners are meant to beat.

With state shared and locked, the next things worth adding are a policy check such as Checkov in front of apply, and separate keys or workspaces per environment so a plan can never be applied against the wrong stack. The Terraform course covers backends, workspaces and CI apply pipelines as one sequence.

Related posts

Quick reference