State, locking, and working as a team
The state file is the tool's memory. Share it remotely, lock it during changes, recover a stale lock carefully, and protect it as a secret.
State is identity
When the tool creates a database, it records in state that module.ledger.aws_db_instance.main is the real object with ID ledger-prod-7f3a. On the next plan it reads that object's current attributes and compares them with the code. State is not a cache that could be rebuilt by looking around. It is the identity mapping between code and reality. Lose it, and the tool believes nothing exists and plans to create everything again next to the originals. Edit it carelessly, and the tool may "forget" a resource it should still manage, or adopt one it should not.
Sharing state: remote backends
A state file on one engineer's laptop works for one engineer. Teams need shared state, and Terraform: Up & Running walks through the requirements:
- Shared storage everyone (and CI) reads and writes.
- Locking so two applies cannot write at once.
- Encryption and access control, because state contains secrets.
- Versioning of the stored file, so a bad write can be rolled back.
Object stores with a lock table, databases, and hosted services all implement this as a remote backend. The configuration names the backend, and every run fetches state, takes the lock, works, writes state, and releases the lock.
When a lock outlives its holder
Locks are released at the end of a run. A run that dies, because its runner was preempted or its laptop lid closed, may leave the lock behind, and every later run fails with "Error acquiring the state lock". The error prints the lock info: ID, who, operation, and when.
force-unlock <ID> releases it, and running it blindly is how teams corrupt state. The recovery is evidence first:
- Read the lock info. Which job, which user, when, and what operation?
- Prove the holder is gone. Check the CI system for that run, check that the runner no longer exists, and ask the person named.
- Unlock exactly that ID.
- Plan before anything else. An interrupted apply may have half-finished. The plan shows what exists now versus what state recorded.
- Fix the cause. Find out why runs can die without cleanup, and give long applies a timeout and a runner that will not be preempted.
The lab follows this sequence. The validator checks that state is intact and consistent afterwards, not just that the error went away.
Isolate blast radius
One state file for "everything" means one mistake can touch everything. The book compares two isolation approaches. Workspaces keep several states in one backend for the same code, which is convenient but makes it easy to apply to the wrong one. File layout gives each environment (and each major component) its own directory and backend configuration. The second makes the environment explicit in every command and every review, and it is what the book recommends for production.
Apply from automation
The team chapter recommends that applies to shared environments run from a pipeline, not from laptops. The pipeline runs plan on every pull request and posts the output for review, then runs apply of that reviewed plan after merge. That gives one place where credentials live, one log of every change, and the same lock discipline every time. It is the Continuous Delivery track's pipeline thinking applied to infrastructure.
Key terms
- State
- The tool's record mapping each resource address in the code to a real object ID, plus cached attributes.
- Remote backend
- Shared storage for state (an object store, a database, a hosted service), usually with locking.
- State lock
- A lease that stops two runs from writing state at once. Normally it is released automatically at the end of a run.
- force-unlock
- Manually releasing a lock by ID. It is safe only after you have proved that the holder is no longer running.
Read further
- Terraform Up & Running, 3rd edition, Ch. 3, "How to Manage Terraform State" (Purchase)
What state is, why local state fails for teams, remote backends with locking, the limitations of backends, and isolating state through workspaces versus file layout. - Terraform Up & Running, 3rd edition, Ch. 10, "How to Use Terraform as a Team" (Purchase)
The workflow for infrastructure code (review, plan in CI, apply from automation) and why apply should not run from laptops.