CPU pinned since the deploy ALERT-7730

OpenLinux · Medium · Recover · about 25 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Alertmanager opened ALERT-7730 at 09:12SEV-3

depot-1's CPU has been at 100% since the 09:10 Waybill deploy. Someone in chat wants to pkill -f waybill.

Customers are still being served, slowly. Killing everything with waybill in its name would also kill the API.

"Before anyone kills anything: find out what is burning CPU and what keeps starting it. If you kill it and walk away, it will be back in 30 seconds." (Mara)

The search indexer runs from a timer every 30 seconds. The deploy also rendered its config from a template.

Your task

Stop the runaway indexer without touching the Waybill API process, fix the config that makes it spin, and leave scheduled indexing working: one indexer run must complete normally.

On the machine

  • ps, top -b -n1, pgrep -af waybill
  • bin/hostctl list-timers, bin/hostctl status waybill-index
  • etc/waybill/indexer.json
  • var/log/deploy.log

Timeline

09:10Waybill 2.4.3 deployed to depot-1, with a new indexer config.
09:12ALERT-7730: node CPU above 95% for 5 minutes.
09:14#ops-depot: "Anyone mind if I just pkill -f waybill?"
09:15Mara: "Yes, I mind. Find it first."

Done when

  1. No indexer process is burning CPU.
  2. The timer is on again, and a fresh indexer run completes with all 1200 parcels.
  3. The Waybill API keeps its original main PID throughout.

Hints

Hint 1

Identify process ownership.

Hint 2

Check who restarts PID 955.

Hint 3

A timer can recreate a killed process.

Hint 4

Stop timer temporarily, fix job, then re-enable.

Show the solution

Stop `waybill-index.timer` so it cannot respawn the job, stop `waybill-index.service`, restore a positive `batch_size` in `etc/waybill/indexer.json` (the deploy rendered an unset variable as 0), run the service once and read its journal, then start the timer again.