Observable Queue

Make a background worker diagnosable before it pages the team. About 16 hours.

The assignment

Docklight receives parcel events and emits notifications. Add signals that distinguish queue backlog, slow dependencies, and worker failure; then write an alert that maps to user impact.

What you hand in

  • Instrumented service and sample workload
  • Prometheus rules and a focused Grafana dashboard
  • Alert test cases for normal, backlog, and dependency-failure states
  • On-call runbook with evidence-first diagnosis

How it gets reviewed

  1. A reviewer can identify which failure occurred from the supplied signals.
  2. A short transient spike does not page; sustained customer impact does.
  3. Metric labels have bounded cardinality.
  4. Runbook steps reproduce and verify recovery.