Alerting on SLOs
Page on the rate at which the error budget is burning, using multiple windows, so you catch fast outages quickly and slow leaks before they cost the month.
From thresholds to budgets
The Observability and Monitoring track built alerts on hand-chosen thresholds: error ratio above 5% for 20 seconds. Once a service has an SLO, there is a better question to ask: how fast are we spending the error budget? An outage that burns 10% of the month's budget in an hour needs a human now. A slow leak that would use the budget over three weeks needs a ticket. A 30-second blip that costs 0.01% needs nothing.
Burn rate
Burn rate is the speed of budget consumption relative to the SLO. At burn rate 1, the budget lasts exactly the window. At burn rate 10, it lasts a tenth of the window. With a 99.9% SLO the budget is 0.1% errors, so an observed error ratio of 1.44% is a burn rate of 14.4. Over a 30-day window that would exhaust the budget in about two days, and every hour at that rate costs 2% of the month's budget.
The general relation is:
budget consumed = burn rate × (alert window / SLO window)
That lets you choose thresholds in terms of what matters, the share of budget lost, instead of raw percentages.
Measuring an alert
The workbook judges every alerting strategy on four attributes:
- Precision: are the alerts real?
- Recall: are real events alerted on?
- Detection time: how quickly does it fire?
- Reset time: how long does it keep firing after the problem is fixed?
It then works through six approaches in turn: target error rate over short windows, longer windows, alert durations, burn-rate alerts, multiple burn rates, and finally multiwindow, multi-burn-rate alerts. Each fixes a weakness of the one before.
The recommended shape
For a 99.9% SLO over 30 days, the workbook's starting configuration is:
| Severity | Long window | Short window | Burn rate | Budget consumed |
|---|---|---|---|---|
| Page | 1 hour | 5 minutes | 14.4 | 2% |
| Page | 6 hours | 30 minutes | 6 | 5% |
| Ticket | 3 days | 6 hours | 1 | 10% |
An alert fires only when both windows exceed the burn rate. The long window ensures the burn is significant, and the short window ensures it is still happening, so the alert resets within minutes of recovery. In PromQL, using recorded error ratios, the first tier reads:
(
job:slo_errors_per_request:ratio_rate1h{job="waybill"} > (14.4 * 0.001)
and
job:slo_errors_per_request:ratio_rate5m{job="waybill"} > (14.4 * 0.001)
)
The constants come from the SLO, so changing the objective changes every alert consistently.
Practicalities
- Low traffic makes ratios jumpy. One failure in ten requests is a 10% error rate. The workbook suggests synthetic traffic, combining services, or minimum request counts.
- Latency SLOs work the same way, using the proportion of requests slower than the threshold as the "bad" ratio. That builds on the Observability and Monitoring track's histograms.
- Dependencies' SLOs are context, not pages, unless your users feel them.
- Keep a small number of SLO alerts per service. They replace most threshold alerts rather than adding to them.
Key terms
- Burn rate
- How fast the error budget is being consumed relative to the SLO. A burn rate of 1 uses exactly the whole budget over the window.
- Precision
- The fraction of alerts that correspond to significant events. Low precision means noise.
- Recall
- The fraction of significant events that produce an alert. Low recall means missed incidents.
- Multiwindow alert
- An alert that requires a high burn rate over both a long and a short window, so it fires on real burns and resets quickly once they stop.
Read further
- The Site Reliability Workbook, Ch. 5, "Alerting on SLOs" (Free to read online)
The four alert attributes (precision, recall, detection time, reset time), the six approaches compared in turn, and the recommended multiwindow, multi-burn-rate configuration with its parameters. - Site Reliability Engineering, Ch. 10, "Practical Alerting" (Free to read online (CC BY-NC-ND 4.0))
Review the ideas on alerting from time-series data. Compare its rule style with the burn-rate rules from the workbook.