Processes, signals, and resources
Find the process that is actually responsible, stop it with the least force, and read process states.
Finding the process that matters
When a host is slow, the first job is to name the consumer. top (or htop) sorted by CPU or memory gives candidates. Then establish who each candidate is before acting on it: ps -o pid,ppid,user,stat,etime,args -p <pid>. The full command line and the parent tell you whether it is the service, a helper, a cron job, or someone's interactive session. /proc/<pid>/ holds the rest: cwd, exe, fd/, and status.
The tempting shortcut is to blame the process whose name matches the alert. An alert says "Waybill is slow". The CPU is actually consumed by an indexer started by the same deploy. Measure first, then attribute.
Signals: asking, insisting, pausing
A signal is a small message the kernel delivers to a process. The ones operators use:
| Signal | Number | Default | Catchable | Typical use |
|---|---|---|---|---|
SIGHUP |
1 | exit | yes | By convention, reload configuration |
SIGINT |
2 | exit | yes | Ctrl-C from a terminal |
SIGKILL |
9 | exit | no | Last resort, with no cleanup |
SIGTERM |
15 | exit | yes | Polite shutdown, the default of kill |
SIGSTOP / SIGCONT |
— | pause / resume | no / yes | Freeze a process without ending it |
The order of escalation is TERM, wait, verify, then KILL only if needed. A program that catches SIGTERM flushes buffers, finishes the current item, releases locks, and exits cleanly. SIGKILL gives it no chance to do any of that. Files can be left half-written and locks left behind for the next run to trip over. Service managers follow the same pattern: systemctl stop sends SIGTERM, waits for a timeout, then sends SIGKILL.
After signalling, verify. Check that the PID is gone (kill -0 <pid> fails), the resource has recovered, and the service you meant to protect is still answering.
Reading process states
ps and top show a state letter:
Rrunning or runnable,Sinterruptible sleep (waiting for an event, normal for idle daemons).Duninterruptible sleep. The process is waiting inside the kernel, usually on I/O, and does not respond even toSIGKILLuntil the I/O completes. ManyDprocesses point at storage trouble.Tstopped, bySIGSTOPor a debugger.Zzombie. The process has exited and its parent has not collected the status. It costs a table slot, not memory. Fix the parent.
Load is not CPU
Linux's load average counts tasks that are runnable plus tasks in uninterruptible sleep. A load of 12 on four CPUs can therefore mean CPU saturation or a pile of processes stuck on a slow disk. Confirm with vmstat 1. r is the run queue, b is the blocked count, and us/sy/wa split the CPU time. Systems Performance names this idea the USE method: for each resource, check utilisation, saturation, and errors. It returns in the “Memory, limits, and a method for performance” notes.
Priority, not punishment
nice and renice lower a process's CPU priority without stopping it. It is a good choice for a batch job you cannot kill but must not let starve the service. ionice does the same for disk I/O on schedulers that support it. Both are mitigations. The incident note should still say why the job ran at the wrong time.
Key terms
- Signal
- An asynchronous notification sent to a process. Most can be caught and handled.
SIGKILLandSIGSTOPcannot. - SIGTERM
- Signal 15, a polite request to exit. Programs catch it to flush data and close connections. It is the default for
kill. - Zombie
- A process that has exited but whose parent has not yet collected its status. It uses no CPU or memory, only a process-table slot.
- Load average
- On Linux, the average number of tasks that are runnable or in uninterruptible sleep, over 1, 5, and 15 minutes.
Read further
- UNIX and Linux System Administration Handbook, 5th edition, Ch. 4, "Process Control" (Purchase)
Process attributes (PID, PPID, UID, niceness), the signal table and which signals can be caught, the state codes shown byps, and the sections ontopand/proc. - How Linux Works, 3rd edition, Ch. 8, "A Closer Look at Processes and Resource Utilization" (Purchase)
Measuring CPU time, the load average discussion, and per-process resource monitoring withtop,lsof, andstrace.