Field notes

Habits that come up again and again in the labs. They are short on purpose: print them, stick them next to the monitor, ignore the ones you already do.

Before you change anything

  • Write down the symptom, who is affected, and when it started. Times in UTC.
  • Reproduce it yourself, from where the user is. curl -v beats a dashboard.
  • Name one other cause it could be, and what would rule it out.
  • Copy the file before you edit it: cp nginx.conf nginx.conf.$(date +%F).

While you are changing it

  • One change at a time, then check again. Two changes at once teach you nothing.
  • Prefer the change you can undo. Revert a commit; do not reset a shared branch.
  • Validate config before reloading: nginx -t, sshd -t, visudo -c, promtool check rules.
  • Match processes by PID, not by name. pkill matches more than you think.

After the change

  • Check the path the customer takes, not just localhost.
  • Ask whether the fix survives a reboot, a redeploy, and the next run of automation.
  • Leave the evidence: logs, core dumps, the old config. Someone will ask on Monday.

Writing it up

  • Timeline first: what happened, when, how you knew.
  • Describe conditions, not culprits. "The deploy job did not need the test job" is useful; a name is not.
  • Every action item has an owner and a ticket, or it will not happen.

Patch, the raccoon in the corner, has one rule: write down what you saw before you write what you think happened.