/var is full and nothing is in it INC-2304

OpenLinux · Medium · Investigate · about 25 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Tomas Berg opened INC-2304 at 06:10SEV-2

depot-2 cannot print shipping labels. /var is at 100%, the night shift deleted the 3.9 GB log, and df still says 100%.

depot-2 is the harbour depot; the 06:00 trucks are already waiting for labels. Waybill's health check returns 507 (insufficient storage).

"/var hit 100% around 02:30. trace.log was 3.9G so I deleted it, then compressed the old access logs. df still says 100% and I have no idea why. Leaving this for day shift, sorry." (Tomas)

Engineering also needs the trace lines around the 02:14 panic for the postmortem. Those lines are in the file Tomas deleted. A file is only really gone when nothing has it open any more.

Your task

Save the trace lines around the 02:14 panic under var/incident/ before anything else, then free the space without restarting Waybill blindly, and give the trace a size limit so this cannot happen again.

On the machine

  • bin/df -h for the simulated 4 GiB /var
  • lsof +L1, ls -l /proc/PID/fd: open files that were deleted
  • etc/waybill/trace.json, re-read on bin/hostctl reload waybill
  • home/nightops/.bash_history

Timeline

02:14Waybill panics; a trace burst starts writing to trace.log.
02:30/var reaches 100%. Label printing fails.
02:41Tomas deletes trace.log (3.9 GB) and compresses access logs. df: still 100%.
06:10Tomas hands over: "no idea why, sorry".

Done when

  1. /var has room again.
  2. Waybill's health check returns 200.
  3. The trace has a size limit, and a rotation drill proves it works.
  4. The 02:14 trace lines are saved under var/incident/.

Hints

Hint 1

Compare df and du.

Hint 2

Look for unlinked open files.

Hint 3

A process can hold deleted file blocks.

Hint 4

Inspect `lsof +L1` and rotate/restart safely.

Show the solution

Find the descriptor with `lsof +L1` or `ls -l /proc/PID/fd`, copy the head of the deleted trace (for example `head -c 64k /proc/PID/fd/3 | grep -a -A4 TRACE-7f3a > var/incident/trace-0214.txt`), set a bounded `max_bytes` in `etc/waybill/trace.json`, then `bin/hostctl reload waybill` so Waybill re-reads its config and reopens the trace. Recheck `bin/df -h` and the health endpoint.