A 30 second crash became a four minute outage OPS-2331

Open2 versionsIncident response · Hard · Fix · about 35 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Mara Okafor opened OPS-2331 at 11:00task

Ledger crashed for 30 seconds at 09:40. It came back, and then stayed down for four more minutes under a flood of retries from our own clients.

Two thousand Waybill clients call Ledger. When Ledger crashed, each one retried every 200 milliseconds, forever. When Ledger came back, it faced every request from the outage at once, failed them, and they were retried again.

"Ledger was down for 30 seconds. We were down for four minutes, and the other three and a half were us." (Mara)

Read clients/retry.json and the incident file before assuming what the clients do today: a policy that ends the storm by giving up early fails the depots in a different way.

bin/storm-sim replays the outage against a retry policy and draws the load Ledger saw. It is a model, not a load test, but its arithmetic is the same arithmetic that took Ledger down.

Your task

Change clients/retry.json so that, in the replay, Ledger recovers within a minute of coming back, fewer than 10% of failed requests are abandoned, and retries are bounded.

On the machine

  • incident/NW-330.md
  • clients/retry.json and clients/README.md
  • bin/storm-sim

Timeline

09:40:00Ledger crashes during compaction; its supervisor restarts it.
09:40:30Ledger is back. Error rate stays at 100%.
09:43:10Waybill clients blocked at the firewall. Ledger recovers in seconds.
11:00OPS-2331: fix the retry policy before the next blip.

Done when

  1. In the replay, Ledger recovers within 60 seconds of coming back.
  2. Fewer than 10% of the requests that failed are abandoned.
  3. Retries are bounded: 2 to 12 attempts, and no single wait longer than 60 seconds.

Hints

Hint 1

Run bin/storm-sim before changing anything. What happens to load right after the outage?

Hint 2

A multiplier above 1 makes each wait longer than the last (exponential backoff).

Hint 3

Add up the waits: they need to cover more than the 30 second outage, or requests give up too early.

Hint 4

Something like 8 attempts, 1000 ms base, multiplier 2, 30000 ms cap works. Try jitter settings and compare.

Show the solution

Run bin/storm-sim on the current policy: retrying forever with no backoff keeps Ledger overloaded for minutes, and three quick tries abandon almost every request during the outage. Set max_attempts to about 8 to 10, base_delay_ms 1000, multiplier 2, max_delay_ms 30000, and a jitter setting. The waits grow to cover the 30 second outage, the backlog drains instead of compounding, and Ledger recovers within about half a minute.