Library
Concepts
26Turn a symptom into a tested hypothesis instead of a guess.
Keep config out of code, keep secrets out of history, and assume anything pushed is exposed.
Describe the state each host should be in, let the tool converge it, and keep each setting in exactly one place.
A database connection is a process with memory. One client that never lets go can lock out everyone else.
Each container has its own loopback; peers find each other by service name on a shared network.
A plan is a diff between desired config, recorded state, and reality; read it before you apply it.
Build once, test that exact artifact, deploy that exact artifact, and keep the previous one.
Test persistence across the destructive lifecycle event, and count deliveries, not just survivals.
Git rarely loses committed work immediately; refs move, and objects stay recoverable for a while.
A health signal is only useful if it fails when users would fail.
Names are resolved by a chain of configuration and servers that fails independently of IP connectivity.
Mitigate first, keep a timeline, and learn without blame.
You declare desired state; controllers keep acting until reality matches, which shapes how things fail.
A conflict is two intentions colliding; resolve the meaning, not the markers.
Hosts that should be identical diverge unless versions are declared and checked.
Access is decided along the whole path, and a process should hold only the rights its job needs.
Identify what is running, who owns it, and what will restart it before you send a signal.
Retry transient failures, but back off, add jitter, and give up eventually, so clients do not keep a recovering service down.
A job that started is not a job that succeeded; schedulers run in a minimal, unattended environment.
A supervised process runs in the environment its unit declares, not the one in your shell.
Read before you change, quote what you pass, and send output where you can keep it.
Measure what users experience, set an explicit target, and page on burning the error budget.
A listener answers only on the addresses it bound, and only callers who can route to them reach it.
Bytes, inodes, and still-open deleted files are three different ways to fill a filesystem.
For every resource, check utilisation, saturation, and errors before theorising.
Certificate validation depends on the chain, the name, and the verifier's clock.
Books
28Troubleshooting method, monitoring, SLOs, incidents, and postmortems.
Cited in A troubleshooting method, Deployment pipelines and artifact identity, Durable state and delivery guarantees, Health checks, readiness, and liveness, Incident response and postmortems, Permissions and least privilege, Retries and backoff, Scheduled jobs, SLIs, SLOs, and alerting on symptoms
Worked practice for SLOs, alerting, on-call, incident response, and canaries.
Cited in Configuration and secrets, Configuration management, Incident response and postmortems, SLIs, SLOs, and alerting on symptoms
How Git stores history, and how reset, merge, rewrite, and recovery really work.
Cited in Configuration and secrets, Git history, reset, and recovery, Merging and resolving conflicts
Shell fundamentals, permissions, processes, redirection, and scripting.
Cited in Package versions and configuration drift, Permissions and least privilege, Processes, ownership, and signals, Scheduled jobs, Shell fundamentals for investigation, Storage capacity, open files, and inodes
Broad reference for services, storage, networking, DNS, and access control.
Cited in How name resolution actually happens, Permissions and least privilege, Service managers and unit files, Sockets, bind addresses, and reachability, Storage capacity, open files, and inodes, TLS certificates, trust, and time
Performance methodology (USE), CPU, memory, file systems, and observability tools.
Cited in A troubleshooting method, Processes, ownership, and signals, SLIs, SLOs, and alerting on symptoms, Storage capacity, open files, and inodes, The USE method for resource problems
Stability patterns and antipatterns in production software.
Cited in Health checks, readiness, and liveness, Retries and backoff, Sockets, bind addresses, and reachability
Durability, replication, stream processing, and delivery guarantees.
Deployment pipelines, building binaries once, and release safety.
Flow, feedback, and learning practices across delivery and operations.
Cited in A troubleshooting method, Incident response and postmortems
Delivery performance metrics and the capabilities that predict them.
Config in the environment, disposable processes, backing services, and logs.
Cited in Configuration and secrets, Container networking and service discovery, Durable state and delivery guarantees, Package versions and configuration drift, Service managers and unit files
Pods, labels, services, deployments, and the reconciliation model.
Cited in Container networking and service discovery, Health checks, readiness, and liveness, Kubernetes controllers and reconciliation
State, plans, modules, and safe change of declarative infrastructure.
Cited in Declarative infrastructure and state
Inventories, playbooks, roles, and idempotent configuration.
Cited in Configuration management, Declarative infrastructure and state, Package versions and configuration drift
What the kernel, boot process, user space, devices, and network stack actually do.
Addressing, routing, ICMP, DNS, and TCP connection behaviour seen on the wire.
Latency, TCP and TLS handshakes, and HTTP behaviour that shapes request time.
Sockets, bind, listen, accept, and connect from the programmer's side.
Least privilege, recovery design, safe deployment, and crisis management.
Namespaces, cgroups, capabilities, image contents, and passing secrets to containers.
Monitoring anti-patterns, alert design, statistics, and what to measure at each layer.
Structured events, tracing, high-cardinality debugging, and SLO-based alerting.
The authoritative description of each object, controller, and probe.
Exactly how connections, transactions, backups, and restores behave, from the people who wrote them.
Cited in Connection limits and leaked sessions, Durable state and delivery guarantees
Operating databases the way SREs run services, especially backup, recovery, and change safety.
How AWS decides whether a request is allowed, and how to write least-privilege policies.
Cited in Permissions and least privilege
Bucket policies, ACLs, and Block Public Access, and how they combine.
The notes are our own summaries. To read a book, follow its link to the author or publisher.