SLOs and error budgets
100% is the wrong target. Choose a reliability objective from what users need, measure it, and spend the remainder deliberately.
Reliability is a product decision
The SRE book's chapter on risk starts with a provocation. 100% is the wrong reliability target for almost anything. Users reach a service through phones, Wi-Fi, ISPs, and browsers that are less reliable than the service itself, so they cannot tell 99.99% from 100%. Each extra "nine" costs disproportionately more, in redundancy, slower releases, and engineering time, while buying nothing users notice. The right target is the level at which users are happy, and that is a product decision informed by data.
Indicators, objectives, agreements
Chapter 4 defines the vocabulary:
- SLI (indicator): a measurement of service behaviour. In practice it is best expressed as a ratio of good events / valid events, so that 0% is total failure and 100% is perfect.
- SLO (objective): a target for an SLI over a time window, for example 99.5% of label requests good over a rolling 28 days.
- SLA (agreement): an external contract with consequences. Keep SLOs tighter than SLAs so you react before you owe credits.
The workbook adds a useful distinction. An SLI specification says what you want to measure ("label requests that succeed quickly, as users see them"). An SLI implementation says exactly how ("load balancer logs, non-5xx and < 500 ms, excluding health checks"). Write both. The implementation decides what counts.
Choosing SLIs from user journeys
Start from what users do, not from what is easy to measure. For Northstar:
| Journey | SLI type | Example |
|---|---|---|
| Print a shipping label | Availability + latency | Proportion of /labels requests non-5xx and < 500 ms |
| Track a parcel | Availability | Proportion of /track requests non-5xx |
| Overnight manifest | Freshness / correctness | Proportion of depots whose manifest is complete by 05:00 |
Measure as close to the user as you practically can. Load balancer logs or client telemetry beat application metrics, which cannot see requests that never arrived. The Observability and Monitoring track's black-box probes are another source.
Choosing targets
The book warns against choosing a target by looking at current performance. That locks you into heroics to maintain an accidental level. Choose from user needs, start somewhat loose, and tighten with evidence. Keep the number of SLOs small, a few per service that cover the important journeys.
The error budget
If the SLO is 99.5%, the error budget is the remaining 0.5%: over 10 million requests in the window, 50,000 bad ones are acceptable. The budget changes conversations:
- Releases, experiments, migrations, and incidents all spend from the same budget.
- While budget remains, the team ships. Reliability is good enough by agreement.
- When the budget is exhausted, the error budget policy applies: freeze non-essential releases and prioritise reliability work until the service recovers.
The workbook stresses writing that policy down and getting agreement from product, development, and SRE before it is needed. Then the decision during a bad month is to follow the policy, not to argue.
Budgets are information, not punishment
Spending budget is fine. That is what it is for. Consistently ending windows with most of the budget unspent suggests the team could ship faster, or that the SLO is tighter than users need. Consistently exhausting it suggests the SLO, the system, or the release process needs attention. Review targets periodically, as the workbook recommends.
Key terms
- SLI
- Service level indicator, a measured ratio of good events to valid events, such as successful label requests faster than 500 ms.
- SLO
- Service level objective, a target for an SLI over a window, such as 99.5% over 28 days.
- SLA
- Service level agreement, a contract with consequences (credits, penalties) if a level is missed. It is usually looser than the SLO.
- Error budget
1 − SLO, the unreliability the service may "spend" in the window on releases, experiments, and failures.
Read further
- Site Reliability Engineering, Ch. 3, "Embracing Risk" (Free to read online (CC BY-NC-ND 4.0))
Measuring service risk, risk tolerance for consumer and infrastructure services, and the motivation for error budgets as a negotiation tool between development and SRE. - Site Reliability Engineering, Ch. 4, "Service Level Objectives" (Free to read online (CC BY-NC-ND 4.0))
The definitions of SLI, SLO, and SLA, choosing indicators by service type, and the advice on choosing targets (keep a safety margin, do not pick a target from current performance). - The Site Reliability Workbook, Ch. 2, "Implementing SLOs" (Free to read online)
The worked example of specifying SLIs, the SLI specification versus implementation distinction, the error budget policy document, and continuous improvement of targets.