Docker said healthy for 14 minutes of 500s INC-2342

Open2 versionsContainers · Medium · Investigate · about 30 min ·Linux + Docker

Lab machine

A private Linux machine with Docker Engine. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Mara Okafor opened INC-2342 at 03:40SEV-2

At 03:12 Ledger crashed. For 14 minutes every tracking request failed, and the load balancer kept sending traffic to Waybill because Docker said it was healthy.

Waybill's health check passes whatever happens. A container can be running, and healthy by that definition, while every request it serves fails.

"The load balancer did exactly what it was told. We told it the wrong thing." (Mara)

Waybill has a readiness endpoint that opens a real connection to Ledger. The stack is in the broken state right now: Ledger is stopped, Waybill says healthy.

Your task

Replace the health check with one that probes the readiness endpoint, with an explicit timeout of 10 seconds or less and an interval short enough to notice an outage quickly. Rebuild, then test both directions: stop Ledger and watch Waybill turn unhealthy, start it and watch it recover.

On the machine

  • docker compose ps, docker inspect --format "{{json .State.Health}}"
  • The HEALTHCHECK in Waybill's Dockerfile or compose.yaml
  • /healthz inside the Waybill container

Timeline

03:12Ledger crashes.
03:12–03:26Every tracking request returns 500. docker ps: waybill (healthy).
03:26Ledger restarted by hand. Traffic recovers.
03:40Mara opens INC-2342: "why did nothing notice?"

Done when

  1. Waybill turns unhealthy while Ledger is down.
  2. It turns healthy again after Ledger recovers.
  3. The health check has an explicit timeout of 10 seconds or less.

Hints

Hint 1

Process state is not readiness. What does the health check actually run?

Hint 2

A health check can come from the image's HEALTHCHECK or from compose.yaml, which overrides it. `docker inspect` shows the one in effect.

Hint 3

Read the probe command to the end: anything that can never fail tells the load balancer nothing.

Hint 4

Probe /healthz with a short timeout, and let a failure be a failure.

Show the solution

`docker inspect --format '{{json .Config.Healthcheck}}'` on the Waybill container shows the check in effect. It is either the image's `HEALTHCHECK CMD true`, or a compose.yaml override whose probe ends in `|| exit 0`. Both succeed whatever happens. Replace it with a bounded probe of the readiness endpoint, for example `wget -q -T 2 -O /dev/null http://127.0.0.1:8080/healthz || exit 1` with `interval: 5s`, `timeout: 3s`, `retries: 2`, wherever the effective check is defined. Rebuild with `docker compose up -d --build`, then stop and start Ledger while watching `docker compose ps`.