Killed on a host with 4 GiB free INC-2399

OpenLinux · Hard · Investigate · about 35 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Teo Marin opened INC-2399 at 11:05SEV-2

Carrier batches have stopped leaving depot-6. The Docklight worker restarts every few seconds, on a host the dashboard calls idle.

The host has plenty of free memory. The worker keeps dying anyway.

"The box has tons of free memory. Maybe add swap?" (on-call chat)

The platform rules say every service keeps a memory limit.

Your task

Find which limit kills the worker, then make batches complete continuously: size the batch or the service limit from the measured peak, stay inside the 192 MiB budget, and keep at least 1000 rows per batch.

On the machine

  • bin/hostctl status docklight-worker (memory current and peak)
  • bin/hostlog -u docklight-worker
  • etc/docklight/worker.env, etc/docklight/capacity.md
  • The worker unit file

Timeline

10:30Carrier contract change: batches of at least 1000 rows.
10:31BATCH_ROWS raised to 6000 "to be safe".
10:32Worker starts restarting every few seconds.
11:05INC-2399. No carrier batch has left since 10:31.

Done when

  1. Batches complete continuously without an OOM kill.
  2. The worker keeps a MemoryMax of 192M or less.
  3. Each batch carries at least 1000 rows.

Hints

Hint 1

Distinguish host memory from cgroup memory.

Hint 2

Read the service’s resource controls.

Hint 3

The process crosses its 256 MiB ceiling.

Hint 4

Reduce batch size or justify a higher unit limit.

Show the solution

`bin/hostctl status docklight-worker` shows the peak climbing into MemoryMax=96M before each kill: every row holds about 16 KiB, so 6000 rows need over 100M. Either lower BATCH_ROWS in `etc/docklight/worker.env` to between 1000 and about 4500 and restart, or raise MemoryMax to at most 192M in the unit, `bin/hostctl daemon-reload`, and restart. Then watch `bin/hostlog -u docklight-worker -f` for batches completing without a kill.