Proxies, load balancers, and the request path
Model a request as a chain of hops, test the middle first, and verify the fix from where the customer stands.
Draw the path before you test it
Every customer request crosses a chain of components, each of which can fail on its own:
scanner → DNS → load balancer → reverse proxy → Waybill :8080 → Ledger / database
Write the chain down at the start of an incident, with the address and port of each hop. Half of all proxy incidents become obvious at this point. The proxy's upstream says :8000, and the service's listener table says :8080.
Split the path
With a chain of six components, testing them in order from one end wastes time. Test in the middle instead. A request from the proxy host straight to the application's address and port answers "is the fault in front of or behind this point?". Then split the remaining half. This is the halving habit from the troubleshooting method concept note, and it turns a vague "the site is down" into a named hop within two or three requests.
Two rules make each test meaningful:
- Test from the upstream side of each hop. "Can the proxy reach the app?" must be tested from the proxy host. Your laptop has different routes, DNS, and firewall rules.
- Use the same name, port, and protocol the real hop uses. Check the proxy's configuration rather than assuming.
What proxies tell you
A reverse proxy is itself an excellent witness. Its error log records why an upstream request failed: connection refused, timeout, reset, or an invalid header. Its access log can record upstream address, status, and timing next to the client status. A client-side 502 paired with "connect() failed (111: Connection refused) while connecting to upstream" names both the hop and the failure shape.
What the proxy is running
A proxy reads its configuration at start and on reload, not every time you save a file. After a change, three things can differ: the file you edited, the configuration the proxy would load now (nginx -t and nginx -T print it, merged from every included file), and the configuration the running process loaded. Included files are read in name order, and a later definition can quietly replace an earlier one, so the file you fixed is not always the one that wins.
Apply a change with a test and a reload. A reload starts new workers with the new configuration and lets the old ones finish their requests. The listener stays open. A restart drops every connection the proxy is carrying, for every site behind it, not just the one you are fixing.
Health checks and load balancers
Load balancers send traffic only to backends whose health check passes, and they remove backends that fail. The SRE book's datacenter load-balancing chapter shows how much depends on that signal. Two failure modes recur:
- The check is too shallow. It requests
/or a static page, so a backend with a dead database stays in rotation and fails every real request. - The check is too deep. It calls every dependency, so one dependency outage removes all backends at once, and partial service becomes total outage.
A good readiness check answers "can this instance serve its main requests right now?" and is cheap enough to run every few seconds. The SRE book also describes a lame duck state: a backend that is shutting down stops accepting new work but finishes what it has. Graceful shutdown in containers and Kubernetes implements the same idea.
Balancing policy
Round robin sends equal request counts to each backend. When requests differ in cost or backends differ in capacity, equal counts produce unequal load. The SRE book discusses policies that use backend feedback, such as least-loaded and weighted round robin, and the subtle failures each one has. For operations work, know which policy your balancer uses and whether its view of backend health matches reality.
Close with the customer's view
After a change, verify from the outside, with the public name, through every hop, from a network like the customer's. Then confirm in the logs that the hop you changed is the one now succeeding. The incident timeline should include each test you ran, its result, and the time. That is the evidence that the halving method is meant to produce.
Key terms
- Reverse proxy
- A server that accepts client requests and forwards them to one or more upstream servers, for example nginx, HAProxy, or Envoy.
- Upstream
- The backend a proxy forwards to, defined by an address and port in its configuration.
- Health check
- A periodic probe a load balancer uses to decide whether to send traffic to a backend.
- Split the path
- Testing the middle hop first so that one result excludes half of the remaining components.
Read further
- Site Reliability Engineering, Ch. 19, "Load Balancing at the Frontend", and Ch. 20, "Load Balancing in the Datacenter" (Free to read online (CC BY-NC-ND 4.0))
In Ch. 19, DNS-based balancing and its limits. In Ch. 20, health checking (the "lame duck" state), and why simple round robin misbehaves with uneven backends. - High Performance Browser Networking, Ch. 1, "Primer on Latency and Bandwidth" (Free to read online)
The components of latency (propagation, transmission, processing, queuing), and why adding hops adds latency that bandwidth cannot remove. - UNIX and Linux System Administration Handbook, 5th edition, Ch. 19, "Web Hosting" (Purchase)
The sections on load balancers and proxies, including reverse proxy configuration and health checks.