It works when Ivo runs it INC-2291

OpenNo account needed3 versionsLinux · Medium · Fix · about 25 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Night shift lead, depot-3 opened INC-2291 at 06:05SEV-3

depot-3 was patched and rebooted at 05:40. Waybill has not come back under the service manager, and the fix Ivo made by hand lasted twenty minutes.

depot-3 is the cold-storage depot. Pallets that are not scanned in within an hour have to be re-checked by hand, so every minute the scanners queue offline costs real work on the dock.

"It works when I run it. Must be something flaky with the service manager." (Ivo, 06:22)

The host runs a small service manager (bin/hostctl, behaves like systemctl) with a journal (bin/hostlog, like journalctl). Ivo's shell history is still in his home directory. A program that runs fine from a shell and fails under a service manager is usually missing something the shell gave it for free.

Your task

Find out why Waybill fails under the service manager but not from Ivo's shell, fix the unit (or the application's config path), reload the unit definitions, and make sure the service manager, not a hand-started copy, is what answers on the Waybill port after a reboot.

On the machine

  • bin/hostctl status waybill, bin/hostlog -u waybill -b
  • etc/units/waybill.service
  • home/ivo/.bash_history: what Ivo ran, and from which directory
  • Health endpoint in TICKET.md

Timeline

05:40depot-3 rebooted for kernel patching.
05:52Handheld scanners: "tracking API unreachable". Cold-storage pallets queue offline.
06:20Ivo starts Waybill by hand, sees 200 on /healthz, logs out.
06:41Scanners fail again.
06:45INC-2291 lands with you.

Done when

  1. waybill.service is enabled and restarts cleanly under the service manager.
  2. /healthz answers 200 after a reboot.
  3. The process answering is the service manager's main PID, not a copy started by hand.

Hints

Hint 1

`bin/hostctl status waybill` first: is the unit failing, or is it not started at all?

Hint 2

The journal line just before the failure says what the server could not find, and where it looked.

Hint 3

Ivo's shell gave the server things the service manager does not. Read his history and his ~/.profile.

Hint 4

Write what the server needs into the unit (WorkingDirectory= or Environment=), daemon-reload, and make sure the unit is enabled.

Show the solution

Compare the managed start with Ivo's. `bin/hostctl status waybill` and `bin/hostlog -u waybill -b` show either a failed start ("unable to open ./config/prod.json", with the directory it ran in) or a unit that is not enabled. Ivo's history and `home/ivo/.profile` show what his shell gave the server: a working directory, or WAYBILL_CONFIG. Put that into the unit (`WorkingDirectory=` with the absolute path of opt/waybill, or `Environment=WAYBILL_CONFIG=<absolute path>`) and `bin/hostctl daemon-reload`, or `bin/hostctl enable waybill` if that was all that was missing. Restart it, stop any copy started by hand, and confirm /healthz after `bin/hostctl reboot`.