Library

Short notes on the ideas behind the labs, each pointing to the labs that use it and the book chapters that cover it properly.

Concepts

26
A troubleshooting method

Turn a symptom into a tested hypothesis instead of a guess.

9 labs
Configuration and secrets

Keep config out of code, keep secrets out of history, and assume anything pushed is exposed.

5 labs
Configuration management

Describe the state each host should be in, let the tool converge it, and keep each setting in exactly one place.

4 labs
Connection limits and leaked sessions

A database connection is a process with memory. One client that never lets go can lock out everyone else.

1 lab
Container networking and service discovery

Each container has its own loopback; peers find each other by service name on a shared network.

2 labs
Declarative infrastructure and state

A plan is a diff between desired config, recorded state, and reality; read it before you apply it.

3 labs
Deployment pipelines and artifact identity

Build once, test that exact artifact, deploy that exact artifact, and keep the previous one.

5 labs
Durable state and delivery guarantees

Test persistence across the destructive lifecycle event, and count deliveries, not just survivals.

3 labs
Git history, reset, and recovery

Git rarely loses committed work immediately; refs move, and objects stay recoverable for a while.

5 labs
Health checks, readiness, and liveness

A health signal is only useful if it fails when users would fail.

4 labs
How name resolution actually happens

Names are resolved by a chain of configuration and servers that fails independently of IP connectivity.

1 lab
Incident response and postmortems

Mitigate first, keep a timeline, and learn without blame.

6 labs
Kubernetes controllers and reconciliation

You declare desired state; controllers keep acting until reality matches, which shapes how things fail.

7 labs
Merging and resolving conflicts

A conflict is two intentions colliding; resolve the meaning, not the markers.

1 lab
Package versions and configuration drift

Hosts that should be identical diverge unless versions are declared and checked.

3 labs
Permissions and least privilege

Access is decided along the whole path, and a process should hold only the rights its job needs.

11 labs
Processes, ownership, and signals

Identify what is running, who owns it, and what will restart it before you send a signal.

3 labs
Retries and backoff

Retry transient failures, but back off, add jitter, and give up eventually, so clients do not keep a recovering service down.

1 lab
Scheduled jobs

A job that started is not a job that succeeded; schedulers run in a minimal, unattended environment.

3 labs
Service managers and unit files

A supervised process runs in the environment its unit declares, not the one in your shell.

6 labs
Shell fundamentals for investigation

Read before you change, quote what you pass, and send output where you can keep it.

10 labs
SLIs, SLOs, and alerting on symptoms

Measure what users experience, set an explicit target, and page on burning the error budget.

5 labs
Sockets, bind addresses, and reachability

A listener answers only on the addresses it bound, and only callers who can route to them reach it.

4 labs
Storage capacity, open files, and inodes

Bytes, inodes, and still-open deleted files are three different ways to fill a filesystem.

3 labs
The USE method for resource problems

For every resource, check utilisation, saturation, and errors before theorising.

5 labs
TLS certificates, trust, and time

Certificate validation depends on the chain, the name, and the verifier's clock.

2 labs

Books

28
Site Reliability EngineeringBetsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy (eds.) · O'Reilly, 2016Free to read online (CC BY-NC-ND 4.0)

Troubleshooting method, monitoring, SLOs, incidents, and postmortems.

Cited in A troubleshooting method, Deployment pipelines and artifact identity, Durable state and delivery guarantees, Health checks, readiness, and liveness, Incident response and postmortems, Permissions and least privilege, Retries and backoff, Scheduled jobs, SLIs, SLOs, and alerting on symptoms

The Site Reliability WorkbookBetsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, Stephen Thorne (eds.) · O'Reilly, 2018Free to read online

Worked practice for SLOs, alerting, on-call, incident response, and canaries.

Cited in Configuration and secrets, Configuration management, Incident response and postmortems, SLIs, SLOs, and alerting on symptoms

Pro Git, 2nd editionScott Chacon, Ben Straub · Apress, 2014Free to read online (CC BY-NC-SA 3.0)

How Git stores history, and how reset, merge, rewrite, and recovery really work.

Cited in Configuration and secrets, Git history, reset, and recovery, Merging and resolving conflicts

The Linux Command LineWilliam Shotts · Seventh Internet EditionFree PDF (CC BY-NC-ND)

Shell fundamentals, permissions, processes, redirection, and scripting.

Cited in Package versions and configuration drift, Permissions and least privilege, Processes, ownership, and signals, Scheduled jobs, Shell fundamentals for investigation, Storage capacity, open files, and inodes

UNIX and Linux System Administration Handbook, 5th editionEvi Nemeth, Garth Snyder, Trent R. Hein, Ben Whaley, Dan Mackin · Addison-Wesley, 2017Purchase

Broad reference for services, storage, networking, DNS, and access control.

Cited in How name resolution actually happens, Permissions and least privilege, Service managers and unit files, Sockets, bind addresses, and reachability, Storage capacity, open files, and inodes, TLS certificates, trust, and time

Systems Performance, 2nd editionBrendan Gregg · Addison-Wesley, 2020Purchase; the USE method summary is free on the author's site

Performance methodology (USE), CPU, memory, file systems, and observability tools.

Cited in A troubleshooting method, Processes, ownership, and signals, SLIs, SLOs, and alerting on symptoms, Storage capacity, open files, and inodes, The USE method for resource problems

Release It!, 2nd editionMichael T. Nygard · Pragmatic Bookshelf, 2018Purchase

Stability patterns and antipatterns in production software.

Cited in Health checks, readiness, and liveness, Retries and backoff, Sockets, bind addresses, and reachability

Designing Data-Intensive ApplicationsMartin Kleppmann · O'Reilly, 2017Purchase

Durability, replication, stream processing, and delivery guarantees.

Cited in Durable state and delivery guarantees

Continuous DeliveryJez Humble, David Farley · Addison-Wesley, 2010Purchase; summaries are free on the book's site

Deployment pipelines, building binaries once, and release safety.

Cited in Deployment pipelines and artifact identity

The DevOps Handbook, 2nd editionGene Kim, Jez Humble, Patrick Debois, John Willis, Nicole Forsgren · IT Revolution, 2021Purchase

Flow, feedback, and learning practices across delivery and operations.

Cited in A troubleshooting method, Incident response and postmortems

AccelerateNicole Forsgren, Jez Humble, Gene Kim · IT Revolution, 2018Purchase; DORA publishes the capability research free

Delivery performance metrics and the capabilities that predict them.

Cited in Deployment pipelines and artifact identity

Kubernetes Up & Running, 3rd editionBrendan Burns, Joe Beda, Kelsey Hightower, Lachlan Evenson · O'Reilly, 2022Purchase

Pods, labels, services, deployments, and the reconciliation model.

Cited in Container networking and service discovery, Health checks, readiness, and liveness, Kubernetes controllers and reconciliation

Terraform Up & Running, 3rd editionYevgeniy Brikman · O'Reilly, 2022Purchase

State, plans, modules, and safe change of declarative infrastructure.

Cited in Declarative infrastructure and state

Ansible for DevOpsJeff Geerling · LeanpubPurchase

Inventories, playbooks, roles, and idempotent configuration.

Cited in Configuration management, Declarative infrastructure and state, Package versions and configuration drift

How Linux Works, 3rd editionBrian Ward · No Starch Press, 2021Purchase

What the kernel, boot process, user space, devices, and network stack actually do.

TCP/IP Illustrated, Volume 1, 2nd editionKevin R. Fall, W. Richard Stevens · Addison-Wesley, 2011Purchase

Addressing, routing, ICMP, DNS, and TCP connection behaviour seen on the wire.

High Performance Browser NetworkingIlya Grigorik · O'Reilly, 2013Free to read online

Latency, TCP and TLS handshakes, and HTTP behaviour that shapes request time.

Beej's Guide to Network ProgrammingBrian "Beej Jorgensen" Hall · Web, continuously revisedFree to read online

Sockets, bind, listen, accept, and connect from the programmer's side.

Building Secure and Reliable SystemsHeather Adkins, Betsy Beyer, Paul Blankinship, Piotr Lewandowski, Ana Oprea, Adam Stubblefield · O'Reilly, 2020Free to read online

Least privilege, recovery design, safe deployment, and crisis management.

Container SecurityLiz Rice · O'Reilly, 2020 (a second edition is available)Purchase

Namespaces, cgroups, capabilities, image contents, and passing secrets to containers.

Practical MonitoringMike Julian · O'Reilly, 2017Purchase

Monitoring anti-patterns, alert design, statistics, and what to measure at each layer.

Observability EngineeringCharity Majors, Liz Fong-Jones, George Miranda · O'Reilly, 2022 (a second edition is available)Purchase; Honeycomb offers a sponsored download

Structured events, tracing, high-cardinality debugging, and SLO-based alerting.

Kubernetes Documentation, ConceptsThe Kubernetes project · Web, versioned with each releaseFree online (CC BY 4.0)

The authoritative description of each object, controller, and probe.

PostgreSQL Documentation, version 17The PostgreSQL Global Development Group · Web, versioned with each releaseFree online (PostgreSQL Licence)

Exactly how connections, transactions, backups, and restores behave, from the people who wrote them.

Cited in Connection limits and leaked sessions, Durable state and delivery guarantees

Database Reliability EngineeringLaine Campbell, Charity Majors · O'Reilly, 2017Paid (book or O'Reilly subscription)

Operating databases the way SREs run services, especially backup, recovery, and change safety.

AWS Identity and Access Management User GuideAmazon Web Services · Web, continuously updatedFree online

How AWS decides whether a request is allowed, and how to write least-privilege policies.

Cited in Permissions and least privilege

Amazon S3 User GuideAmazon Web Services · Web, continuously updatedFree online

Bucket policies, ACLs, and Block Public Access, and how they combine.

The notes are our own summaries. To read a book, follow its link to the author or publisher.