SLIs, SLOs, and alerting on symptoms

Measure what users experience, set an explicit target, and page on burning the error budget.

A service level indicator (SLI) is a measured ratio of good events to valid events. Examples are successful requests over all requests, or requests served under 300 ms over all requests. A service level objective (SLO) is the target for that ratio over a window, such as 99.5% over 28 days. The shortfall allowed by the target is the error budget. Spending it on releases and experiments is fine. Exhausting it means slowing down.

Pick SLIs from the user's side of the system. The SRE book's four golden signals are latency, traffic, errors, and saturation. RED (rate, errors, duration) suits request-driven services, and USE (utilisation, saturation, errors) suits resources. Page on symptoms users feel, such as error ratio or latency. Keep causes like CPU high or disk 80% on dashboards and tickets unless they predict imminent user impact.

Good alerts are:

  • Actionable. Someone can do something now.
  • Tuned by burn rate. Page when the budget is burning fast enough to be exhausted soon, for example by alerting on the error ratio over a short and a long window together. The workbook's alerting chapter compares approaches.
  • Tested. Break the thing and watch the alert fire. Fix it and watch the alert resolve. An alert that has never fired is an assumption.

In Prometheus terms, SLIs are usually expressions such as sum(rate(http_requests_total{code=~"5.."}[5m])) / sum(rate(http_requests_total[5m])), and alerts are rules over those ratios with a for: duration to ride out blips.