Networks drop packets, services restart, and leaders fail over. Retrying a failed request is often the right response, because the second attempt usually succeeds. Retrying badly is one of the most common ways a short failure becomes a long outage.
Picture two thousand clients calling a database that crashes for thirty seconds. If every client retries immediately and keeps retrying, the database comes back to a queue of every request from the outage plus every new one, all arriving at once. It cannot serve them, so they fail, so they are retried. The service is healthy and still down, held under by its own clients. This is sometimes called a retry storm or a thundering herd.
A good retry policy has four parts:
- Backoff. Wait longer after each failure, usually doubling the wait (exponential backoff). This spreads the backlog over time so the recovering service can drain it.
- A cap on each wait. Without one, exponential growth quickly produces waits of minutes or hours, which help nobody.
- A limit on attempts. Some failures are not transient. A bounded number of attempts lets the caller report an error instead of hammering a dead dependency forever.
- Jitter. Add randomness to each wait. When many clients fail at the same instant, identical backoff schedules make them retry in lockstep; jitter breaks up the waves.
The waits should add up to more than the outages you want to survive. A policy that gives up after two seconds protects the server but loses every request during a thirty second restart.
Retries interact with other stability patterns. Timeouts decide when an attempt has failed. Circuit breakers stop calling a dependency that is clearly down and test it occasionally. Load shedding on the server rejects excess work cheaply so it can keep serving some requests. Together they let a system bend under failure instead of collapsing.