The average says 90 milliseconds CX-311

Open2 versionsMonitoring · Hard · Investigate · about 40 min ·Linux + Docker

Lab machine

A private Linux machine with Docker Engine. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Ana Costa opened CX-311 at 11:30SEV-4

Since release 2.4.3, depot-3 staff say some label prints freeze for two seconds. The latency dashboard shows /labels at 90 ms.

If 95 requests take 20 ms and 5 take 1.5 seconds, the average is under 100 ms and five people waited a second and a half. The dashboard plots a mean, and nothing that pages looks at the tail correctly.

"Averages say it's fine. Probably their Wi-Fi." (on-call)

Waybill exports latency as a histogram, so the real distribution is there to be read.

Your task

Record each route's 99th percentile as route:http_request_duration_seconds:p99 and add a page that fires when /labels p99 goes above 0.5 seconds for a short while, and resolves afterwards.

On the machine

  • http_request_duration_seconds_bucket{route,le}
  • prometheus/rules/
  • The current recording rule route:http_request_duration_seconds:mean1m

Timeline

MonRelease 2.4.3.
Tuedepot-3: "hazmat labels hang sometimes".
11:30CX-311. On-call: "Averages say it's fine. Probably their Wi-Fi."

Done when

  1. route:http_request_duration_seconds:p99 matches the histogram's 99th percentile for each route.
  2. A severity page alert fires within about a minute of a slow tail on /labels and resolves after.
  3. It stays quiet while /labels is healthy.

Hints

Hint 1

Compare the mean with the bucket counts above 1 second.

Hint 2

Read every latency rule: what does each one record, and does it return a value for /labels?

Hint 3

histogram_quantile needs bucket rates summed by le, plus the labels you group by. Without le it has nothing to compute.

Hint 4

Record a correct p99 by route, then alert on the /labels series with a short for.

Show the solution

Query the histogram buckets for /labels: a few percent of requests take about two seconds, which the mean flattens into 90 ms. Read `prometheus/rules/`: either the only rule is a mean, or a p99 rule exists but sums the buckets by `route` only, without `le`, so `histogram_quantile` returns nothing and its alert never fires. Record `histogram_quantile(0.99, sum by (route, le) (rate(http_request_duration_seconds_bucket{job="waybill"}[1m])))` as `route:http_request_duration_seconds:p99`, and alert on `route:http_request_duration_seconds:p99{route="/labels"} > 0.5` with `for: 20s` and `severity: page`. Reload Prometheus.