Troubleshooting is a loop, not an inspiration. You start from a report: what the user saw, when, and where. You examine the system's telemetry and state. You diagnose by forming a hypothesis that would explain every symptom. Then you test it by predicting what you will see if you are right, and checking. When you have tested it, you treat the problem and verify the fix from the place the user stands.
Two habits separate calm operators from frantic ones.
Split the path in half. A request from a depot scanner crosses DNS, a proxy, a socket, a process, and a dependency. Test the middle of that chain first. If the proxy can reach the upstream directly, the fault is on the customer side of the proxy, and you have ruled out half the system with one request. The networking labs are built so that this halving works.
Prefer evidence that could prove you wrong. "The process is running" rarely disproves anything. "A request through the proxy returns 502 while a direct request returns 200" does. Before you change anything, write down what result would make you abandon your hypothesis.
Common traps:
- Fixing the first anomaly you see. Production systems are full of harmless oddities. Ask whether the anomaly explains this symptom and its start time.
- Verifying from the wrong side. A loopback health check proves only a local path. Verify from the user's side of every hop you changed.
- Changing several things at once. You learn nothing about which change mattered, and you cannot roll back cleanly.
The Systems Performance methodologies chapter catalogues named approaches. Among them are the USE method, drill-down analysis, and the "streetlight" anti-method of looking only where it is convenient. Recognising which one you are using keeps an investigation honest.