No space left, 41% used INC-2317

Open3 versionsLinux · Medium · Investigate · about 25 min ·Linux

Lab machine

A private Linux machine with the problem already set up. Sessions last up to 60 minutes.
Depot-4 shift lead opened INC-2317 at 13:20SEV-2

The depot-4 spooler rejects every new manifest with "No space left on device". df says the spool is 41% full.

depot-4 is the spool hub: every label printed in the north region passes through its spooler. Two restarts changed nothing.

"Spooler says No space left on device but df -h shows the spool at 41%. I restarted it twice. Is the disk lying?" (depot-4 shift lead)

A filesystem can run out of two things: bytes, and inodes, the entries that describe each file. A few million tiny files use almost no bytes and every inode. The retention job has reported success every hour for a month.

Your task

Find out what is using the inodes and why the hourly retention job never removes it, fix the job, and run it once. Keep customer manifests in spool/outbound/ and any retry marker younger than seven days.

On the machine

  • bin/df -h and bin/df -i for the simulated /spool
  • opt/spooler/CHANGELOG.md
  • var/log/spool-cleanup.log
  • The retention unit and its find command under etc/units/

Timeline

Jul 30Spooler 3.2.0 changes how it names retry markers.
Every hourThe retention job runs, finds nothing to delete, and logs success.
13:02Spooler: "No space left on device".
13:20INC-2317: outbound manifests backing up.

Done when

  1. The spool has free inodes again.
  2. The spooler is writing manifests again.
  3. The hourly retention job now removes expired markers, and only those.

Hints

Hint 1

`bin/df -i`: a filesystem can run out of files before it runs out of bytes.

Hint 2

Count files per directory. The growth is thousands of tiny files, not one big one.

Hint 3

`var/log/spool-cleanup.log` says the hourly job removes nothing, or stopped running. Since when, and what changed then?

Hint 4

Compare the job's find command and its timer with what `opt/spooler/CHANGELOG.md` says about markers. Fix the job, then run it once.

Show the solution

`bin/df -i` shows the spool out of inodes while bytes are fine, and `spool/retry/` holds thousands of markers. Find out why the hourly job stopped expiring them: compare `opt/spooler/bin/cleanup-retry-markers` with where and how spooler 3.2 writes markers (`opt/spooler/CHANGELOG.md`), and check `bin/hostctl list-timers` and `var/log/spool-cleanup.log` to see whether it runs at all. Fix that (the name pattern, the depth it searches, or the timer), run `bin/hostctl start spool-cleanup.service` once, confirm with `bin/df -i`, and watch the spooler recover in its journal.