Incident response and postmortems

Mitigate first, keep a timeline, and learn without blame.

During an incident, the goal is to restore service, not to find the root cause. Those goals often conflict. Rolling back a suspect deploy restores service in minutes, while understanding why it failed can take a day. Prefer reversible mitigations, and preserve evidence as you go.

Structure keeps a response calm:

  • Roles. One incident commander coordinates, others operate, and one person communicates. A person operating should not also be managing stakeholders.
  • A live timeline. Write down timestamps for what you observed, what you changed, and what happened next. This becomes the postmortem's spine, and it prevents two people making conflicting changes.
  • Handovers. State what is known, what was tried, what is running, and what to watch. The night on-call's shell history in the disk-full lab is a handover that failed.

A postmortem records impact, timeline, detection, triggers, contributing conditions, and action items with owners. The SRE book's framing is that it should be blameless. Asking "why did the system allow this?" produces fixes like alerts, guardrails, and automation. Asking "who did this?" produces silence.

Distinguish the trigger (the deploy, the reboot, the rotation) from contributing conditions (no health check, no retention, untested rollback). Trigger-only postmortems fix the one path that failed and leave the others open.