What to monitor
Symptoms before causes, the four golden signals, and the anti-patterns that make monitoring noisy and blind at once.
Two questions monitoring must answer
The SRE book frames monitoring around two questions: what is broken, and why. The first is answered by symptoms, what users experience: errors, slowness, unavailability. The second is answered by causes, such as a full disk, an exhausted connection pool, or a failed dependency. Both matter, but they serve different purposes:
- Symptoms drive alerts. A user-facing problem exists, and someone must act.
- Causes drive dashboards and investigation. Here is where to look.
An alerting setup built on causes ("page if memory > 90%") fails twice. It pages for conditions that hurt nobody, and it stays silent for failures whose cause nobody predicted. The errors-without-a-page lab is the second failure exactly.
The four golden signals
For a user-facing service, the book recommends measuring at least:
- Latency: how long requests take. Measure successes and failures separately, because a fast error is still an error.
- Traffic: demand on the system, such as requests per second, by route.
- Errors: the rate of failed requests, explicit (5xx) and implicit (wrong content, too slow to count).
- Saturation: how "full" the service is, meaning the resource closest to its limit, often a leading indicator.
These describe the service from its users' side. The Linux Administration track's USE method describes each resource from the inside. You need both, the golden signals to notice and USE to explain.
White box and black box
White-box monitoring uses internals the service exposes: metrics, logs, and traces. It is detailed and often points at causes. Black-box monitoring tests behaviour from outside, as a user would. A probe requests the public URL, through DNS, TLS, and the load balancer. It sees what white-box cannot, such as an expired certificate, a broken route, or a proxy pointing at the wrong port, all from the Networking for Operators track. Use black-box probes for the few user journeys that matter most and white-box metrics for everything else.
Anti-patterns
Practical Monitoring opens with the mistakes most teams make and pairs each with a design pattern:
- Tool obsession: believing the next tool will fix monitoring. Decide what to measure first.
- Monitoring as a job: a separate team owns monitoring, so the service owners do not. Service owners should instrument and own their alerts.
- Checkbox monitoring: collecting easy metrics (CPU, ping) so that "we have monitoring", without knowing whether the service works. Start from the user's perspective.
- Using monitoring as a crutch: piling alerts on a fragile system instead of fixing it.
- Manual configuration: targets and alerts added by hand drift out of date. Configure monitoring automatically from the same source as deployment.
Monitor the monitoring
An alert over metrics that have stopped arriving evaluates to nothing, and nothing does not fire. If a deploy moves the metrics port, renames a job, or breaks the exporter, every alert on that service goes silent at the moment you most need it. Two safeguards are needed: an alert on scrape health (up{job="waybill"} == 0), and an alert on absence (absent(up{job="waybill"}), which covers the case where the target disappeared from discovery entirely). Then test them. Break the scrape on purpose in staging and confirm the page arrives.
Key terms
- Symptom
- What users experience, such as errors, slowness, or unavailability. "Is it broken?"
- Cause
- A reason for a symptom, such as a full disk, high memory, or a failed dependency. "Why is it broken?"
- Golden signals
- Latency, traffic, errors, and saturation. The SRE book's minimum set for a user-facing system.
- Black-box monitoring
- Testing externally visible behaviour as a user would, for example a probe that requests the public URL.
Read further
- Site Reliability Engineering, Ch. 6, "Monitoring Distributed Systems" (Free to read online (CC BY-NC-ND 4.0))
Definitions (white-box, black-box, dashboards, alerts), why to monitor, symptoms versus causes, the four golden signals, worrying about your tail, and choosing an appropriate resolution for measurements. - Practical Monitoring, Ch. 1, "Monitoring Anti-Patterns", and Ch. 2, "Monitoring Design Patterns" (Purchase)
Each anti-pattern and its matching design pattern, especially composable monitoring, monitoring from the user's perspective, and continual improvement.