The fix passed its tests and did not ship INC-2445

Open2 versionsCI/CD · Hard · Incident · about 55 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
PagerDuty opened INC-2445 at 14:32SEV-2

Paged at 14:32: hazmat checkouts failing since the 14:20 release. That release's pipeline was green and fixed the very test that failed at 13:40.

This one is a real incident with three phases, in this order: stop the bleeding, fix the cause, write it up.

"Customers first. Get production back on something known good, by digest. Then figure out how a green pipeline shipped the old code." (Mara)

There is a break-glass release token in ci/secrets.json for exactly this.

Your task

Roll production back to the last known-good image by digest, fix the workflow so it deploys exactly what it tested, and fill in POSTMORTEM.md: impact, a timeline with the bad and good digests and the commit the bad image was really built from, contributing conditions, and action items, without blame.

On the machine

  • bin/prod status, bin/deployctl history, bin/ci runs
  • repo/.github/workflows/release.yml
  • POSTMORTEM.md template

Timeline

13:40Pipeline fails: hazmat checkout test.
14:05Fix committed; tests pass.
14:20Release deploys. Hazmat checkouts start failing.
14:32You are paged.

Done when

  1. Production is back on the last known-good image, deployed by digest, with no checkout errors.
  2. The pipeline deploys exactly the commit it tested, and failing tests still block.
  3. POSTMORTEM.md covers impact, a timeline with both digests and the real source commit, contributing conditions, and action items, without blame.

Hints

Hint 1

Mitigate customer impact first. `bin/deployctl history` lists the digest that ran before the incident.

Hint 2

Deploy it by digest: `RELEASE_TOKEN=... bin/deployctl deploy waybill@sha256:...`. The token is in `ci/secrets.json`.

Hint 3

Compare the tested commit in `bin/ci runs` with the running image's revision label and its `BUILD` file. Then find how the build got other code: a cache, or a checked-out ref.

Hint 4

Build after the tests from the checked-out commit, pass the digest to the deploy job as a job output, and deploy `waybill@<digest>`. Then fill in `POSTMORTEM.md`.

Show the solution

Roll back by digest to the image that ran before the incident and confirm 0% errors. Then compare the bad image's revision label with the commit the 14:20 run tested: the image was built from the 13:40 commit, either restored from a `dist/` cache keyed only on `VERSION`, or checked out from `release/candidate`, a branch cut at 13:40. Make the build job need the test job, build from the pushed commit with no cache or pinned ref, tag per commit, and deploy the job's digest output. The postmortem's timeline names the bad and known-good digests and the commit the bad image was really built from. Its contributing conditions are how the stale code reached the build, a build job that ran without passing tests, and a deploy by the mutable `latest` tag.