Incident response
Declare early, assign roles, mitigate before you fully understand, and keep a living record of what is happening.
Incidents need structure
The SRE book's incident chapter opens with an unmanaged incident. A capable engineer tries to fix everything alone, colleagues freelance changes without telling anyone, communication with stakeholders lapses, and the outage drags on. Nobody did anything foolish. Nobody was coordinating. The fix is a lightweight structure adapted from emergency services' Incident Command System, and both books describe it.
Roles
Separate the responsibilities, even when one person holds several of them in a small incident:
- Incident commander (IC): owns the incident. Keeps the overall picture, assigns roles, sets priorities, and decides. The IC does not debug.
- Operations lead: works on the system and is the only role making production changes. Others propose, and operations executes and records.
- Communications lead: sends regular updates to stakeholders and users, so engineers are not interrupted by questions.
- Planning lead: handles longer-running needs, such as tracking follow-ups, arranging handoffs, and ordering food at hour six.
Roles scale recursively. In a large incident, sub-teams get their own leads reporting to the IC.
Declare early
The book's guidance is to declare an incident if any of these hold: you need to involve a second team, the outage is visible to customers, or the issue is unresolved after an hour of focused analysis. Declaring costs little. Standing down an incident that turned out minor is easy, and organising one late, after hours of uncoordinated work, is not.
Mitigate first
Users care that service is restored, not that the cause is understood. During an incident, prefer actions that stop the harm:
- Roll back a recent deploy or configuration change (the Continuous Delivery track makes this routine).
- Fail over, drain a bad zone, or shift traffic.
- Shed load or disable a feature with a toggle.
Rolling back a suspect change is often right even before it is proven to be the cause. Mitigation is reversible, and the investigation continues with less pressure. Preserve evidence as you go: keep the bad artifact, the logs, and the timeline.
The live document
A shared incident document, updated continuously, is the command post's memory:
- Status: impact, scope, current severity.
- Timeline: timestamped observations, decisions, and actions, in UTC.
- Hypotheses: what is being tested, by whom, and with what result.
- Actions: who is doing what, and what is pending.
- Communications log: what was told to whom, and when.
It makes handoffs possible, keeps late joiners from asking repeated questions, and becomes the backbone of the postmortem.
Troubleshooting under pressure
The troubleshooting loop from the SRE book's chapter 12 (report, triage, examine, diagnose, test, treat) and the path-splitting habit from the Networking for Operators track still apply, but stress makes people skip steps. Watch for the pitfalls the chapter names: fixating on symptoms that are irrelevant, misunderstanding how the system works, chasing coincidences, and favouring a recent change as the cause without testing it. Write each hypothesis in the document with the result that would disprove it.
Clean handoffs
Incidents outlast shifts. The IC role changes hands explicitly: "You are now IC. Do you accept?" "Yes." The live document is current, and in-flight actions are listed. A sustainable on-call rotation (chapter 11) is part of incident response too. Tired responders make errors, so plan for relief before you need it.
Key terms
- Incident commander (IC)
- Holds the overall picture, assigns roles and tasks, and decides. The IC does not do the hands-on debugging.
- Operations lead
- Works on the system to mitigate, and is the only role changing production during the incident.
- Mitigation
- Restoring service, for example by rollback, failover, or shedding load, before the root cause is fully understood.
- Live incident document
- A shared, continuously updated record of status, timeline, hypotheses, actions, and owners.
Read further
- Site Reliability Engineering, Ch. 14, "Managing Incidents" (Free to read online (CC BY-NC-ND 4.0))
The unmanaged-incident case study, the elements of incident management (recursive separation of responsibilities, a recognised command post, a live incident state document, clear handoff), and when to declare an incident. - The Site Reliability Workbook, Ch. 9, "Incident Response" (Free to read online)
The Incident Command System roles adapted for software, the case studies, and the best practices summary. - Site Reliability Engineering, Ch. 11, "Being On-Call", and Ch. 12, "Effective Troubleshooting" (Free to read online (CC BY-NC-ND 4.0))
In Ch. 11, balanced on-call load and the feeling of safety. In Ch. 12, the report, triage, examine, diagnose, test, and treat loop, and the pitfalls that appear under pressure.