Designing for recovery
Assume something will go wrong. Know the intended state, keep the ability to roll back or revoke, and have a tested path back.
Assume you will need to recover
The BSRS chapter on recovery starts from a premise operators know well. Failures and compromises are certain. What varies is how long recovery takes and how much is lost. Systems can be designed to make recovery fast, safe, and possible at all. Most of the principles come down to three questions:
- Do we know what the system should look like? (intended state)
- Can we get back there? (rollback, restore, rebuild)
- Can we take away what an attacker or a mistake left behind? (revocation)
Know your intended state
To recover, you need an authoritative description of "correct", ideally down to the bytes: which artifact digests run where, what configuration each service has, who has which access, and which data should exist. Infrastructure as code, versioned configuration, release records with provenance, and access-control definitions in code are all intended-state records. Without them, recovery becomes archaeology. The release-history incident in the Version Control with Git track is a small example. Once the tag was gone, only an independent release record could settle what had shipped.
Intended state also enables continuous validation: regularly comparing reality to the record and alerting on drift. Compromise and misconfiguration both show up as drift.
Rollback is not free
Reliability engineering likes rollback, and security engineering is wary of it. A rollback restores a known-good version, which may also be a known-vulnerable version. An attacker who can trigger a rollback can reintroduce a fixed bug. BSRS suggests keeping rollback available but bounded, with a minimum acceptable version or a deny list of known-bad releases enforced at the deployment choke point. The same reasoning applies to configuration and certificates.
Data adds its own constraint. Rolling back code is easy only if the data still fits the old code (the Continuous Delivery track's expand and contract). Rolling back data means restoring a backup, which loses everything written since.
Revocation must be designed in
When a credential, key, or certificate is compromised, you need to stop trusting it everywhere, quickly. That works only if:
- every consumer checks revocation (or credentials are short-lived enough to expire soon),
- the revocation mechanism scales and is fast,
- and there is a re-issue path that does not depend on the compromised thing.
Short-lived credentials (the “Managing secrets” notes) are partly a revocation strategy, because waiting out the lifetime is a revocation that always works.
Emergency access, again
Recovery often happens when normal systems are down: the identity provider is unreachable, CI is broken, or the network is partitioned. BSRS insists that emergency access paths exist and are tested, and that they are audited and alerting, which is the “Least privilege by design” notes’ break-glass. Improvised emergency access, such as sharing a root password in chat, causes the next incident.
Rehearse
Every recovery mechanism needs regular exercise: restoring a backup into a scratch environment, rolling back a release, rotating a credential, rebuilding a host from code. Measure time to recover and data lost, and compare them with what the business needs. A recovery plan that has never run is a guess. The lab capstones across this academy are rehearsals of this kind, done on purpose and checked by a validator.
Key terms
- Intended state
- An authoritative description of what a system should be, such as versions, configuration, and access, against which recovery is measured.
- Revocation
- Invalidating a credential, certificate, or artifact so it can no longer be used, quickly and verifiably.
- Minimum acceptable version
- A floor below which rollbacks are refused, so a known-vulnerable version cannot be redeployed.
- Recovery rehearsal
- Regularly exercising restore, rollback, and revocation procedures so they work when needed.
Read further
- Building Secure and Reliable Systems, Ch. 9, "Design for Recovery" (Free to read online)
The design principles for recovery. Read closely the sections on rollbacks as a trade-off between security and reliability, explicit revocation mechanisms, knowing your intended state, and designing for testing and continuous validation. Note the guidance on emergency access. - Building Secure and Reliable Systems, Ch. 8, "Design for Resilience" (Free to read online)
Defence in depth, controlling degradation, and blast-radius controls (compartments and failure domains). - Building Secure and Reliable Systems, Ch. 18, "Recovery and Aftermath" (Free to read online)
Planning recovery after an incident, including the scope of recovery and the postmortem. It links the Secure Operations track to the Site Reliability Engineering track.