Alerting that pages for the right things
Every page should be urgent, actionable, and about user impact. Everything else goes to a ticket or a dashboard.
What a page is for
A page interrupts a person, at dinner or at 03:00, and asks for judgement. The SRE book holds pages to a strict standard. Every paging rule should be:
- Urgent: user impact is happening or imminent.
- Actionable: there is something a human should do now.
- About symptoms: users are, or will be, affected. "A server is unhappy" is not enough.
- Requiring intelligence: if the response is a script, automate the script instead.
Everything that fails these tests still has a place. Tickets cover work that matters but can wait, such as slow disk growth, certificate expiry in 21 days, or a degraded replica. Dashboards hold context for investigation, such as CPU, memory, and per-dependency latency.
Ratios, not counts
Alert on what users experience, normalised by traffic. An error count pages at peak because more traffic means more absolute errors, and misses failures at night. An error ratio does not:
- alert: WaybillRouteErrorRatioHigh
expr: |
sum by (route) (rate(http_requests_total{job="waybill",code=~"5.."}[1m]))
/
sum by (route) (rate(http_requests_total{job="waybill"}[1m]))
> 0.05
for: 20s
labels: { severity: page }
annotations:
summary: '{{ $labels.route }} is failing {{ $value | humanizePercentage }} of requests'
sum by (route) keeps the dimension responders need. Without it, one broken route among ten healthy ones averages below the threshold.
Speed versus noise
Two knobs balance detection speed against false alarms:
- Rate window (
[1m],[5m]): longer windows smooth noise and react more slowly. forduration: how long the condition must hold. It filters blips at the cost of delay.
Choose both from the impact. A total outage of checkout must page within a minute or two. A 2% error rate may be fine to tolerate for longer. The Site Reliability Engineering track replaces hand-tuned thresholds with burn-rate alerts derived from an SLO, which is the systematic answer to this trade-off.
When there is no data
An alert expression over series that no longer exist returns an empty result, and an empty result never fires. A broken exporter, a moved metrics port, or a renamed job silences every alert for that service, as in the metrics-went-dark lab. Every service with paging alerts needs:
up{job="waybill"} == 0 or absent(up{job="waybill"})
The first part catches a target that is discovered but failing to scrape. The second catches a target that has vanished from discovery entirely.
Keep alerts honest over time
Practical Monitoring recommends reviewing alerts regularly against what responders actually did. For each alert that fired, was it real, did someone act, and could the action be automated? Alerts acknowledged and ignored week after week should be deleted, retuned, or turned into tickets. Every page should link to a runbook that explains what it means and the first three things to check. If nobody can write that runbook, the alert is not ready to page.
Key terms
- Page
- An alert that interrupts a human immediately. Reserve it for urgent, actionable user impact.
- Ticket
- An alert that creates work for business hours. It suits slow-burning or non-urgent problems.
- Error ratio
- Failed requests divided by all requests over a window. It is independent of traffic volume, unlike an error count.
- for duration
- How long a condition must hold before the alert fires. It suppresses blips at the cost of detection delay.
Read further
- Site Reliability Engineering, Ch. 10, "Practical Alerting" (Free to read online (CC BY-NC-ND 4.0))
Alerting on time-series data, rules and aggregation, alerting on rates and ratios rather than counts, and routing to pages versus tickets. - Site Reliability Engineering, Ch. 6, "Monitoring Distributed Systems" (Free to read online (CC BY-NC-ND 4.0))
Re-read "Tying These Principles Together" and "Monitoring for the Long Term". Note the questions every new paging rule must answer. - Practical Monitoring, Ch. 3, "Alerts, On-Call, and Incident Management" (Purchase)
What makes a good alert, stopping alert fatigue, and the advice to delete or tune alerts that do not lead to action.