"The process is up" and "the service can do its job" are different claims. Load balancers, orchestrators, and humans act on health signals, so a signal that always says "healthy" is worse than none. It routes customers to a broken instance with full confidence.
Distinguish three questions:
- Liveness: is the process stuck beyond recovery, so that restarting it would help? Keep this check cheap and local. Failing liveness on a dependency outage restarts every replica at once and makes things worse.
- Readiness: should this instance receive traffic right now? This check may include critical dependencies, such as the database connection, because pulling a replica out of rotation is safe and reversible.
- Startup: has initialisation finished? This keeps slow starts from tripping the other two.
Make the probe itself safe:
- Bound it. Set an explicit timeout shorter than the interval. A hung dependency must not hang the check.
- Tune detection time. The time to notice an outage is roughly interval × retries. Defaults are often minutes.
- Check the path users take. An HTTP request to the real endpoint beats
pgrep, and a check that exercises the dependency beats one that returns a constant.
Verify a health check the way you verify any alert: break the dependency deliberately and time how long the signal takes to turn. Then restore it and confirm recovery. The SRE book's monitoring chapter makes the broader point. A signal is only worth keeping if it leads to the right action, whether a person reads it or a machine does.