Friday evening, everything at once NW-360

OpenCapstone · Hard · Boss incident · about 60 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Ana Costa opened NW-360 at 19:42SEV-1

Waybill 2.6 went out at 16:40. Since 19:30 every depot's tracking page shows errors. The engineer before you rebooted the host once and went home sick.

It is Friday, 19:42. Every depot's tracking page shows errors, support has 40 calls waiting, and the on-call engineer who was on this went home sick after rebooting the host once.

"I do not need to know everything yet. I need to know when customers can track parcels again, and then I need to know everything." (Ana)

More than one thing is wrong, and some of them hide the others. Keep incident/timeline.md as you go: Monday's review will read it, and you will not remember the order by then.

Your task

Restore tracking for customers through the edge, make the whole host come back on its own after a reboot, stop /var filling up, run the missed backup and turn its timer back on, and record each fault and fix in incident/timeline.md.

On the machine

  • TICKET.md and ops-urls.txt
  • bin/hostctl list-units, bin/hostctl status <unit>, bin/hostlog -u <unit>
  • bin/df-var
  • opt/waybill/CHANGELOG.md
  • etc/units/, etc/waybill/waybill.env, etc/edge/edge.conf

Timeline

ThuLoad test LT-77 adds a rate limit at the edge. "TODO remove after the test."
Fri 16:40Waybill 2.6 deployed. Its port setting is renamed.
Fri 19:05The on-call engineer reboots the host once, then goes home sick.
Fri 19:30Tracking errors for every depot. /var keeps filling.
Fri 19:42NW-360, SEV-1.

Done when

  1. Customers reach Waybill through the edge, even under normal traffic.
  2. Everything comes back on its own after a reboot.
  3. /var is below 50% full, and debug logging is off so it stays that way.
  4. The backup timer is enabled and a fresh backup with the Ledger journal exists.
  5. incident/timeline.md names each fault and what you changed.

Hints

Hint 1

Start from the customer URL in ops-urls.txt and follow the request inward. Try it more than once.

Hint 2

bin/hostctl list-units shows what is running. Is Waybill one of them, and is it enabled?

Hint 3

Start Waybill, read bin/hostlog -u waybill, then opt/waybill/CHANGELOG.md.

Hint 4

bin/df-var, then var/log/waybill: what setting keeps writing there? The backup refuses to run until /var has room.

Show the solution

Rename PORT to WAYBILL_PORT in etc/waybill/waybill.env and set LOG_LEVEL=info; empty var/log/waybill/debug.log; enable and start waybill.service; remove the load test's rate_limit_per_minute from etc/edge/edge.conf; enable backup.timer and start backup.service once; write one line per fault in incident/timeline.md.