Health checks, readiness, and liveness

A health signal is only useful if it fails when users would fail.

"The process is up" and "the service can do its job" are different claims. Load balancers, orchestrators, and humans act on health signals, so a signal that always says "healthy" is worse than none. It routes customers to a broken instance with full confidence.

Distinguish three questions:

  • Liveness: is the process stuck beyond recovery, so that restarting it would help? Keep this check cheap and local. Failing liveness on a dependency outage restarts every replica at once and makes things worse.
  • Readiness: should this instance receive traffic right now? This check may include critical dependencies, such as the database connection, because pulling a replica out of rotation is safe and reversible.
  • Startup: has initialisation finished? This keeps slow starts from tripping the other two.

Make the probe itself safe:

  • Bound it. Set an explicit timeout shorter than the interval. A hung dependency must not hang the check.
  • Tune detection time. The time to notice an outage is roughly interval × retries. Defaults are often minutes.
  • Check the path users take. An HTTP request to the real endpoint beats pgrep, and a check that exercises the dependency beats one that returns a constant.

Verify a health check the way you verify any alert: break the dependency deliberately and time how long the signal takes to turn. Then restore it and confirm recovery. The SRE book's monitoring chapter makes the broader point. A signal is only worth keeping if it leads to the right action, whether a person reads it or a machine does.