The timeout change never took effect OPS-2320

Open2 versionsConfiguration management · Medium · Fix · about 30 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Teo Marin opened OPS-2320 at 16:00SEV-4

CHG-5520 raised the request timeout from 5 to 30 seconds. Every waybill.conf says 30s. The running services still time out after 5.

A service reads its config when it starts. The playbook has a handler that restarts Waybill when the config changes. Somewhere between the change and the restart, the link is cut.

"We quietened that playbook during the 2.3 migration because it was too noisy. I did not realise it would stop the restarts too." (Teo)

Fixing the playbook is half of it. The files already say 30s, so the next run will see nothing to change.

Your task

Make the template task report changes honestly so the handler restarts Waybill when, and only when, the config changes. Then get every running Waybill onto the current config once.

On the machine

  • site.yml (tasks and handlers)
  • bin/svc status depot-N shows what each running process loaded
  • group_vars/depots.yml

Timeline

AugCHG-5402 adds changed_when: false to the template task "to quiet the nightly report".
MonCHG-5520: request_timeout 30s. The playbook runs; files change; nothing restarts.
16:00Carriers still time out after 5 s. OPS-2320.

Done when

  1. Every running Waybill uses the request_timeout from group_vars.
  2. When the config changes, one playbook run restarts Waybill on every host.
  3. A run without config changes restarts nothing, and the template task reports changes honestly.

Hints

Hint 1

bin/svc status depot-1 shows what the running process loaded.

Hint 2

A handler runs only when a task that notifies it reports changed, and only if its own conditions allow. Run the playbook and read what it says about handlers.

Hint 3

Look at both ends of the link: the template task (does it report changes?) and the handler (does it have a `when:`?).

Hint 4

The files already say 30s, so the next run will not notify anything. Restart each host once with bin/svc restart, or run the playbook with a one-off restart.

Show the solution

`bin/svc status` shows each process still running with 5s. Read both ends of the notify link in `site.yml`: either the template task has `changed_when: false`, so it never notifies, or the handler has a `when:` on a flag nobody turned back on, so it is skipped. Remove the one that cuts the link. The files already changed last week, so restart each host once (`bin/svc restart depot-1`, `depot-2`, `depot-3`) and confirm with `bin/svc status`.