Locked out by a job that no longer exists DEP-844

Open2 versionsInfrastructure as code · Hard · Incident · about 35 min ·Linux + Docker

Lab machine

A private Linux machine with Docker Engine. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Dmitri Vos opened DEP-844 at 08:30SEV-4

The hazmat flag for depot-7 never shipped. An apply job stopped halfway, and every run since says the state is locked. Someone suggests -lock=false.

A state lock stops two runs from writing the same state at once. If the holder dies, the lock stays. Removing it is safe only once you know for sure the holder is gone.

"Just run it with -lock=false, it's only one container." (release channel)

The CI records say what happened to the job that holds the lock, and to its runner. Read them before you touch the lock.

Your task

Confirm the lock's holder is dead, release the lock through OpenTofu, apply the pending hazmat change, and leave state consistent with what is running.

On the machine

  • The lock info in the error message
  • ci/jobs/4411.json, ci/runners.json
  • bin/tofu -chdir=infra
  • The state service (bin/hostctl status state-server)

Timeline

Tue 18:02Job #4411 starts applying FEATURE_HAZMAT=on.
Tue 18:03CI runner 17 is preempted. The job dies holding the lock.
WedEvery plan: "Error acquiring the state lock".
08:30"Just run it with -lock=false."

Done when

  1. The holder of the lock is confirmed dead before its lock is removed.
  2. The lock is released through OpenTofu, not by editing backend storage.
  3. FEATURE_HAZMAT=on is live, recorded in state, and the stack plans clean.

Hints

Hint 1

The lock info says who holds it and since when.

Hint 2

Find that holder in the CI records: which job, which runner, and whether either still exists. Not every job in ci/jobs is the one you want.

Hint 3

`tofu force-unlock <ID>` releases a lock through the backend.

Hint 4

If a write was refused and OpenTofu saved errored.tfstate, push it after unlocking, then plan.

Show the solution

Read the lock info from `bin/tofu -chdir=infra plan`: its ID and its holder, `runner@<runner>`. Match it in `ci/jobs/` and `ci/runners.json`: the apply job that took it failed when its spot runner was reclaimed, or was cancelled and its runner recycled, and the runner is deprovisioned. Job 4415 on the other stack is alive and irrelevant. Then `bin/tofu -chdir=infra force-unlock <ID>`, `bin/tofu -chdir=infra apply`, and confirm a clean plan. If you already tried `-lock=false`, the backend refused the write and left `errored.tfstate`; unlock, run `tofu state push errored.tfstate`, then plan.