The alert that could not see PM-2417

Open2 versionsMonitoring · Hard · Incident · about 35 min ·Linux + Docker

Lab machine

A private Linux machine with Docker Engine. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Mara Okafor opened PM-2417 at 10:40SEV-3

After the Waybill 2.4.4 deploy, labels failed for one in three requests for eleven minutes. The error alert exists, promtool says it is valid, and nobody was paged.

"We added the alert after the last postmortem. It is right there. Why didn't it fire?" (Mara)

Your task

Restore scraping, and add a page that fires within a minute whenever Waybill's metrics disappear, for any reason, and resolves when they return. No page while Waybill is healthy.

On the machine

  • Prometheus UI: Status → Targets
  • prometheus/ config and rules
  • CHANGELOG-2.4.4.md

Timeline

09:10Waybill 2.4.4 deployed. Its metrics move to a new port (CHANGELOG-2.4.4.md).
09:12–09:231 in 3 /labels requests fail. No page.
10:40PM-2417: "paged for nothing, again".

Done when

  1. Prometheus scrapes Waybill successfully.
  2. A severity page alert fires within about a minute when Waybill's metrics vanish, and resolves when they return.
  3. No page fires while Waybill is healthy.

Hints

Hint 1

Status → Targets shows each scrape's health and last error. A healthy target can still bring in less than you think.

Hint 2

Query the metric the alert uses. If it returns nothing, find where it is lost: the port, or the scrape's relabelling.

Hint 3

Rules over absent series produce nothing to alert on.

Hint 4

Restore the data, then alert on up == 0 (and absent up) with a short for.

Show the solution

Query `http_requests_total{job="waybill"}`: it returns nothing, so the error alert had nothing to evaluate. Find why in Status → Targets and `prometheus/prometheus.yml`: either the scrape still targets port 8080 while 2.4.4 serves metrics on 9464 (`CHANGELOG-2.4.4.md`), or it targets 9464 but a `metric_relabel_configs` keep-list written for the 2.3 metric names drops the request metrics. Fix that, then add an alert on `up{job="waybill"} == 0 or absent(up{job="waybill"})` with `for: 20s` and `severity: page`, reload Prometheus, and confirm the series and the target.