22 minutes of failed labels, nobody paged PM-2402

Open2 versionsMonitoring · Medium · Fix · about 35 min ·Linux + Docker

Lab machine

A private Linux machine with Docker Engine. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Mara Okafor opened PM-2402 at 09:00task

Yesterday from 14:02 to 14:24 one in five label requests failed. Depot staff phoned it in. Monitoring stayed green, and nobody was paged.

The postmortem had one action item for you: alert on the symptom customers felt, a share of failing requests, not on a cause such as memory.

"Page when people cannot print labels. Not when one request out of two hundred fails, and not on counts that change with traffic." (Mara)

Prometheus is running with Waybill and steady depot traffic. You can inject errors and watch.

Your task

Add a rule that pages within about a minute when any route's error ratio passes 5%, stays quiet under normal background errors, resolves after recovery, and passes promtool.

On the machine

  • Prometheus UI: docker compose port prometheus 9090
  • prometheus/ rules (reload via the /-/reload endpoint)
  • http_requests_total{route,code}

Timeline

14:02/labels starts returning 500 to about 20% of requests.
14:15depot-1 phones support.
14:24Recovers on its own.
09:00Postmortem action item: page on what customers feel.

Done when

  1. A severity page alert fires within about a minute of a 20% error burst on /labels.
  2. It stays quiet under normal traffic and resolves after recovery.
  3. promtool accepts the configuration.

Hints

Hint 1

Explore the counters in the Prometheus UI before writing or fixing a rule.

Hint 2

Read every rule that pages: what it watches, and how long its condition must hold.

Hint 3

Divide errors by all requests, per route, and compare the ratio with a threshold.

Hint 4

A short for (seconds, not half an hour) keeps one failed scrape from paging, without sitting out a real outage.

Show the solution

Query `rate(http_requests_total{code=~"5.."}[1m])` by route: the errors were visible all along. Read `prometheus/rules/`: either the only paging rule watches memory, or an error-ratio rule exists but must hold for 30 minutes, longer than the whole 22-minute burst. Page on `sum by (route) (rate(http_requests_total{job="waybill",code=~"5.."}[1m])) / sum by (route) (rate(http_requests_total{job="waybill"}[1m])) > 0.05` with `for: 20s` and `severity: page`, then reload Prometheus. Keep the memory alert as it is.