Stability patterns and overload
Integration points fail, and the system must not fail with them. Timeouts, circuit breakers, bulkheads, back pressure, and load shedding.
Integration points are where systems die
Release It! starts from production experience: the most dangerous parts of a system are its integration points, every call to another service, database, or API. Remote calls can fail fast, fail slowly, hang, return garbage, or succeed and take forever. The book's stability antipatterns describe how one bad integration point spreads:
- Blocked threads: callers wait on a slow dependency until every worker is stuck.
- Chain reactions: one instance fails, its load shifts to its peers, and they fail in turn.
- Cascading failures: a failure in one layer propagates to the layers calling it.
- Slow responses: worse than fast failures, because they tie up resources on both sides.
- Unbounded result sets: a query that returned 100 rows in testing returns ten million in production.
The common thread is a small problem in one place consuming a shared resource (threads, connections, memory, capacity) that everything else needs.
Patterns that contain failure
Each stability pattern counters one or more antipatterns:
Timeouts. Every remote call needs one, sized to the caller's own deadline. Without timeouts, every other pattern fails, because a hung call holds its thread forever. Propagate deadlines downstream (the SRE book recommends it) so work abandoned by the client is not still being done three services deep.
Retries, carefully. Retrying a transient failure is fine. Retrying immediately, many times, at every layer multiplies load during exactly the moment the system is weakest. Use exponential back-off with jitter, a small maximum, a retry budget (for example, at most 10% extra requests), and retries at one layer only.
Circuit breaker. Track failures to a dependency. When they cross a threshold, open the circuit and fail calls immediately. After a cool-off, let a trial request through (half-open). If it succeeds, close the circuit. The caller saves its resources, and the dependency gets breathing room. Combine with a fallback, such as cached data or a degraded feature, where one exists.
Bulkheads. Partition resources so a failure stays local: separate pools for Ledger calls and for label rendering, separate instances for critical and batch traffic. A flood in one compartment does not sink the ship.
Fail fast. If a request cannot succeed, for example because a dependency's circuit is open or a required resource is missing, reject it at once instead of doing partial work.
Steady state. Anything that accumulates (logs, temp files, sessions, caches, queues) needs an automatic bound or purge. Otherwise the system eventually fails from its own debris, as in the Linux Administration track's notes on disk-full and inode incidents.
Overload and back pressure
The SRE book's chapters on overload and cascading failures add the capacity view. As load approaches capacity, latency climbs, requests time out, clients retry, and the effective load rises further, a positive feedback loop that can take down a healthy fleet. The defences:
- Load shedding: when overloaded, reject excess requests early and cheaply, so the requests you accept still finish within deadline. Throughput of successful requests (goodput) stays high.
- Back pressure: bounded queues that push "slow down" upstream instead of buffering indefinitely.
- Criticality: tag requests by importance (label printing above analytics) and shed the least critical first.
- Client-side throttling: clients that see many rejections reduce their own request rate.
- Graceful degradation: serve a reduced but useful response, such as tracking without the map, rather than none.
Health checks are part of stability
The Containers in Production track and the Kubernetes Operations track warned against health checks that fail on dependency problems. In stability terms, such checks cause chain reactions. Every instance is marked unhealthy, restarted, or removed at once, and a partial outage becomes total. A good design checks what the instance itself controls, reports dependency trouble through readiness or degraded responses, and lets circuit breakers handle the dependency.
Reviewing a service
For each dependency of a service, ask: Is there a timeout, and is it shorter than our own deadline? What is the retry policy? Is there a circuit breaker and a fallback? Is the connection pool isolated? What happens to our users when this dependency is slow rather than down? Writing the answers down for Waybill, covering Ledger, the label renderer, and the database, is a worthwhile exercise. It is also the most effective reliability review most teams never do.
Key terms
- Timeout
- A bound on how long to wait for a remote call. Without one, a slow dependency holds your resources indefinitely.
- Circuit breaker
- Stops calling a failing dependency after repeated errors, fails fast for a while, then probes to see if it has recovered.
- Bulkhead
- Partitioning resources (thread pools, connection pools, instances) so a failure in one part cannot exhaust all of them.
- Load shedding
- Deliberately rejecting some work when overloaded, so the rest can succeed.
Read further
- Release It!, 2nd edition, Ch. 4, "Stability Antipatterns", and Ch. 5, "Stability Patterns" (Purchase)
Integration points, chain reactions, cascading failures, blocked threads, slow responses, and unbounded result sets. Then timeouts, circuit breaker, bulkheads, steady state, fail fast, shed load, and create back pressure. For each pattern, note which antipattern it counters. - Site Reliability Engineering, Ch. 22, "Addressing Cascading Failures" (Free to read online (CC BY-NC-ND 4.0))
Causes of cascading failure (server overload, resource exhaustion, retries), preventing overload, and the guidance on retries and deadlines. - Site Reliability Engineering, Ch. 21, "Handling Overload" (Free to read online (CC BY-NC-ND 4.0))
Per-customer limits, client-side throttling, criticality, and why handling overload gracefully beats trying never to be overloaded.