Processes, ownership, and signals

Identify what is running, who owns it, and what will restart it before you send a signal.

Every process has a PID, a parent, an owner, a command line, and resource counters. Before you signal anything, establish four facts:

  1. Which process is actually consuming the resource? Use ps -o pid,ppid,user,pcpu,rss,etime,cmd or top. Do not trust the name of the service that was blamed in chat.
  2. Who started it? Follow the parent PID, or ask the service manager which unit owns the process. A child of a timer-driven unit comes back on the timer's schedule.
  3. What does it serve? Killing a process that shares a name pattern with the customer API (pkill -f waybill) turns a slow service into an outage.
  4. What will stop it permanently? Usually that means disabling a schedule or fixing an input, not repeating kill.

Signals are messages, not just weapons. SIGTERM asks a process to exit cleanly, and a service manager sends it first on stop. SIGKILL cannot be caught, so it leaves no chance to flush state. SIGHUP conventionally means "reload configuration and reopen files". A process sees exactly one signal; anything that restarts it afterwards is policy outside the process.

CPU percentages have a unit. On Linux, ps reports percent of one CPU, so a busy single-threaded loop shows about 100% and a busy two-thread process about 200%. A runaway loop that makes no system calls shows high user time and no I/O wait. That pattern points at logic and input, such as a batch size of zero that never advances an offset, rather than at the disk or the network.

A kill is a mitigation: it buys time. The remediation changes the cause, then proves one bounded run succeeds before the schedule is restored.