Logs, time, and scheduled jobs
Reconstruct events across restarts and rotations, and make scheduled work run the way it runs by hand.
Two log systems, one question
Most distributions run the systemd journal, often alongside a traditional syslog daemon that writes text files under /var/log. The journal stores structured entries. Every line carries fields such as _SYSTEMD_UNIT, _PID, PRIORITY, and _BOOT_ID. That makes the questions you ask during an incident direct to express:
journalctl -u docklight -b -1 --since "02:00" -p warning
journalctl -u docklight -o short-precise --no-pager | tail -50
journalctl --list-boots
One setting surprises people. If /var/log/journal does not exist (or Storage= is volatile), the journal lives in memory and previous boots are lost. On any host you may need to investigate after a crash, make storage persistent.
Time is evidence
Every log line is only as good as its timestamp. Check three things before building a timeline. Is the clock synchronised (timedatectl)? Which timezone do the logs use? Do different hosts agree? A five-minute skew between a proxy and an application turns "the proxy erred first" into a false story. Prefer UTC for servers and in incident notes.
Rotation is a cooperation
Log files are rotated so they do not grow forever: renamed, compressed, and eventually deleted. A process writes to an open file descriptor, not to a name. After app.log is renamed to app.log.1, the process keeps writing to app.log.1 until it reopens. logrotate offers two answers:
create+postrotate: rename, create a fresh file, then signal the program (oftenSIGHUP) to reopen. This is clean and needs the program's cooperation.copytruncate: copy, then empty the original in place. It works with programs that never reopen, and it can lose the lines written between copy and truncate.
When a log "looks clean", ask whether you are reading the file the process is actually writing. ls -l /proc/<pid>/fd | grep log tells you. Rotation without a reopen is also the usual source of the “Storage, filesystems, and inodes” notes’ deleted-but-open disk usage.
Scheduled jobs run in a different world
A cron job is started by the cron daemon, not by you. Its environment is minimal: a short PATH, often /usr/bin:/bin, few variables, the job owner's home as the working directory, and no terminal. Output is mailed if a mailer exists and otherwise lost. Most "works by hand, fails at 06:00" incidents reduce to one of these:
- A command not on cron's
PATH. - A variable your profile set, absent in cron.
- A relative path.
- A
%in the crontab line, which cron turns into a newline. - Output discarded, so nobody saw the error.
The fix pattern: put the logic in a script with absolute paths and explicit variables, call it from the crontab, and append its output and errors to a log you can read.
Missed runs and doubled runs
Classic cron runs a job when the clock matches while the daemon is running. A host that is down at 06:00 skips the 06:00 run, and nothing catches up. systemd timers with Persistent=true record the last run and fire once at boot if one was missed. anacron serves the same role on some systems. The SRE book's cron chapter frames the trade-off well. Some jobs are worse if skipped, such as a daily invoice run, and some are worse if doubled, such as a charge. Decide which kind you have, and design the job to be idempotent so that a late or repeated run is harmless.
Key terms
- Journal
- systemd's structured log store. Entries carry fields such as unit, PID, priority, and boot ID.
- Log rotation
- Periodically renaming, compressing, and deleting log files so they do not grow without bound.
- copytruncate
- A logrotate mode that copies the file and then empties it in place, for programs that cannot reopen logs. It can lose lines written between the copy and the truncate.
- Persistent timer
- A systemd timer with
Persistent=true. It runs a missed activation at the next boot.
Read further
- UNIX and Linux System Administration Handbook, 5th edition, Ch. 10, "Logging" (Purchase)
The systemd journal andjournalctl, how syslog and the journal coexist, and the section on log rotation withlogrotate. Note the options that control what happens to the file a process has open. - How Linux Works, 3rd edition, Ch. 7, "System Configuration — Logging, System Time, Batch Jobs, and Users" (Purchase)
The sections on system logging, time and clock synchronisation, and scheduling recurring tasks with cron and systemd timer units. - Site Reliability Engineering, Ch. 24, "Distributed Periodic Scheduling with Cron" (Free to read online (CC BY-NC-ND 4.0))
Read the opening sections on cron's reliability model, especially skipped versus doubled launches and idempotency. Skim the distributed design that follows.