The pager cries wolf OPS-2350

OpenIncident response · Medium · Tune · about 35 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Teo Marin opened OPS-2350 at Mon 08:30task

Last week the on-call engineer was paged over two hundred times. Two of those pages were real incidents. She has asked to come off the rotation.

Every alert currently pages, the moment its condition is true for a single minute. The hourly batch job alone pins the CPU 168 times a week.

"By Thursday I had stopped looking at the pager. On Friday the real one came and I nearly missed it." (Teo)

bin/replay runs your rules against last week's real metrics and shows how many times each one would have woken someone, and whether the two real incidents would have paged.

Your task

Tune alerts/rules.yaml so last week would have produced at most four pages and both real incidents still page within 10 minutes. Keep the CPU and disk alerts as tickets, not deleted.

On the machine

  • docs/last-week.md
  • alerts/rules.yaml
  • bin/replay
  • var/metrics/last-week.csv (one row per minute)

Timeline

Last week221 pages. Two were real.
WedNW-341: label API errors for 25 minutes.
FriNW-347: slow label rendering for 40 minutes.
Mon 08:30Teo asks to come off the rotation.

Done when

  1. Both real incidents page within 10 minutes of starting.
  2. The week produces no more than 4 pages.
  3. CPU and disk alerts are kept as tickets, not deleted.

Hints

Hint 1

Run bin/replay. Which rule pages the most?

Hint 2

for_minutes makes a condition last before it fires; blips of 1 to 3 minutes stop paging.

Hint 3

Raise thresholds to levels customers notice, such as errors above 5%.

Hint 4

Change severity to ticket for causes (CPU, disk, queue), keep page for symptoms (errors, latency).

Show the solution

Page on error_ratio_pct above 5 for 5 minutes and p99_ms above 2000 for 5 minutes. Make HighCPU, DiskFilling, and QueueBacklog tickets, with durations long enough to skip routine spikes.