Every process has a PID, a parent, an owner, a command line, and resource counters. Before you signal anything, establish four facts:
- Which process is actually consuming the resource? Use
ps -o pid,ppid,user,pcpu,rss,etime,cmdortop. Do not trust the name of the service that was blamed in chat. - Who started it? Follow the parent PID, or ask the service manager which unit owns the process. A child of a timer-driven unit comes back on the timer's schedule.
- What does it serve? Killing a process that shares a name pattern with the customer API (
pkill -f waybill) turns a slow service into an outage. - What will stop it permanently? Usually that means disabling a schedule or fixing an input, not repeating
kill.
Signals are messages, not just weapons. SIGTERM asks a process to exit cleanly, and a service manager sends it first on stop. SIGKILL cannot be caught, so it leaves no chance to flush state. SIGHUP conventionally means "reload configuration and reopen files". A process sees exactly one signal; anything that restarts it afterwards is policy outside the process.
CPU percentages have a unit. On Linux, ps reports percent of one CPU, so a busy single-threaded loop shows about 100% and a busy two-thread process about 200%. A runaway loop that makes no system calls shows high user time and no I/O wait. That pattern points at logic and input, such as a batch size of zero that never advances an offset, rather than at the disk or the network.
A kill is a mitigation: it buys time. The remediation changes the cause, then proves one bounded run succeeds before the schedule is restored.