Metrics, percentiles, and cardinality
Counters, gauges, and histograms; why averages hide the users who suffer; and how labels multiply cost and power.
Three shapes of metric
Most metrics systems, Prometheus included, offer a few core types:
- Counter: only increases (resetting to zero on restart), for example
http_requests_total. Its raw value is rarely interesting. Its rate is:rate(http_requests_total[1m])is requests per second, and the function handles resets. - Gauge: goes up and down, for example queue depth, memory in use, or open connections. Read the value directly.
- Histogram: counts observations into buckets, for example
http_request_duration_seconds_bucket{le="0.25"}holds the count of requests that took at most 250 ms. Histograms are how you measure distributions such as latency.
Choose the type by the question. "How many requests per second?" needs a counter, "how deep is the queue now?" a gauge, and "how long do requests take?" a histogram.
Averages hide the users who suffer
Practical Monitoring's statistics primer and the SRE book's "worry about your tail" make the same point. The mean is a poor summary of latency. Latency distributions are skewed. Most requests are fast and a few are very slow, and the few are often a specific population, such as hazmat labels, large customers, or one depot. If 97% of label requests take 30 ms and 3% take 2 s, the mean is about 90 ms, and the dashboard looks green while one in thirty-three prints freezes.
Percentiles describe the distribution: p50 (the median), p90, p99. "p99 = 2 s" means one request in a hundred takes 2 s or longer. For user-facing services, alert and set objectives on high percentiles, and watch the median for typical experience.
Computing percentiles correctly
Percentiles cannot be averaged. The average of ten instances' p99 values is not the service's p99, and neither is the maximum. The fix is to aggregate before computing the percentile. Histogram buckets are counts, so they sum correctly across instances, routes, or time:
histogram_quantile(
0.99,
sum by (route, le) (rate(http_request_duration_seconds_bucket{job="waybill"}[1m]))
)
The sum by (route, le) keeps the bucket boundary (le) and the dimension you care about (route) and sums everything else, such as instances and pods. Recording this as a rule (route:http_request_duration_seconds:p99) makes dashboards and alerts cheap and consistent. That is the average-hides-slow-labels lab. Percentiles from histograms are estimates bounded by bucket boundaries, so choose buckets around your objectives (for example, include 0.5 s if 0.5 s is the target).
Resolution matters too
The SRE book also warns about measurement resolution. A one-minute average CPU can hide five-second saturation spikes, and a five-minute rate can hide a 90-second outage. Pick scrape intervals and rate windows to match the speed of the failures you need to see, and remember that longer windows smooth away exactly the short events users notice.
Labels and cardinality
Labels make metrics powerful. sum by (route), by (depot), and by (code) slice one metric many ways. Every unique combination of label values is a separate time series that costs memory, storage, and query time. Labels with a few dozen values (route, status class, region) are fine. Labels with unbounded values (customer ID, request ID, full URL with query string, email) can create millions of series and take down the metrics system.
Observability Engineering argues that this is a structural limit of metrics. They are pre-aggregated, so they cannot answer questions about dimensions you did not anticipate or cannot afford. High-cardinality questions such as "which customer?", "which build?", or "which request?" belong in structured events and traces, which the “Events, logs, and traces” notes cover. Use metrics for the known, bounded questions they answer cheaply.
Key terms
- Counter
- A value that only increases (and resets on restart), such as requests served. Use its rate, not its raw value.
- Gauge
- A value that goes up and down, such as queue length, memory in use, or temperature.
- Histogram
- Counts of observations in buckets (≤ 0.1 s, ≤ 0.25 s, …) from which percentiles can be estimated and correctly aggregated.
- Cardinality
- The number of unique label combinations of a metric. Each is a separate time series with storage and query cost.
Read further
- Practical Monitoring, Ch. 4, "Statistics Primer" (Purchase)
Mean, median, percentiles, and standard deviation. Note why the book warns against averages for latency, and against averaging percentiles. - Site Reliability Engineering, Ch. 6, "Monitoring Distributed Systems" (Free to read online (CC BY-NC-ND 4.0))
Re-read "Worrying About Your Tail (or, Instrumentation and Performance)" and "Choosing an Appropriate Resolution for Measurements". Both are about how you aggregate. - Observability Engineering, Chapter "Structured Events Are the Building Blocks of Observability" (Purchase; Honeycomb offers a sponsored download)
The limitations of metrics as a building block, in particular pre-aggregation and the cost of high-cardinality dimensions.