Releasing safely
Separate deployment from release, expose changes gradually, and make rollback a routine operation.
Deployment is not release
Two words are often used interchangeably, and separating them is the biggest step toward safe change:
- Deployment puts a new version of software into an environment.
- Release makes a capability available to users.
When they are the same event, every deploy is a big-bang exposure to all users. When they are separate, deploys become routine, frequent, and boring, and releases become controlled decisions.
Three techniques
Blue-green deployment. Run two identical production environments. Blue serves traffic while green receives the new version and is verified. Then the router switches traffic to green. Rollback is switching back. The cost is running two environments and handling state, such as databases and sessions, shared between them.
Canary releases. Deploy the new version to a small slice of capacity and route a small share of real traffic to it. Compare its behaviour with the rest. Widen in steps, for example 1%, 10%, 50%, then 100%, or roll back. The SRE Workbook's canarying chapter adds the discipline. Evaluate against a control population running the old version at the same time, because absolute numbers swing with time of day. Pick metrics that reflect user experience (error ratio, latency percentiles), and give each step long enough to see slow failures.
Feature toggles. Ship code paths turned off and enable them by configuration, for internal users first, then a percentage, then everyone. Toggles separate release from deployment entirely. They also carry a cost. Each live toggle doubles the paths to test, so remove them once the feature is fully launched.
Rollback is a deployment
Continuous Delivery treats rollback as deploying the previous known-good artifact with the same automated process. It is not a special emergency procedure. That works only if:
- The previous artifact is still available by digest (the “Build once, and know what you built” notes).
- Configuration for it is versioned and retrievable.
- The database still works with the old version.
- Rollback is rehearsed. A mechanism that runs only in emergencies is untested.
In Git terms, rolling back a change to shared configuration is a revert commit followed by a normal pipeline run (the “Undoing and recovering” notes), not a force-push.
Data changes need two versions at once
During a rolling deploy, canary, or rollback, old and new code run simultaneously against one database. Schema changes must therefore be backward compatible. The standard pattern is expand and contract:
- Expand: add the new column or table. Old code ignores it.
- Migrate: new code writes both shapes, and a backfill copies existing data.
- Switch: reads move to the new shape once it is complete.
- Contract: in a later release, when no running or rollback-target version needs the old shape, remove it.
It is slower than a one-step rename. It is also the difference between a routine rollback and an outage.
Decide in advance what "bad" looks like
Before widening a canary or switching traffic, write down the abort conditions, for example "error ratio above the control by 0.5 percentage points, or p99 latency above 800 ms for five minutes". The Site Reliability Engineering track connects these thresholds to SLOs and error budgets. Without pre-agreed criteria, rollouts continue on optimism.
Key terms
- Blue-green deployment
- Two identical production environments. Traffic switches from the old (blue) to the new (green) at once, and switches back to roll back.
- Canary release
- Sending a small share of traffic to the new version, comparing it with the old, and widening only if it behaves.
- Feature toggle
- A runtime switch that enables code paths independently of deployment.
- Expand and contract
- Changing a schema in backward-compatible steps (add, migrate, then remove) so old and new code can run at the same time.
Read further
- Continuous Delivery, Ch. 10, "Deploying and Releasing Applications" (Purchase; summaries are free on the book's site)
Blue-green deployments, canary releasing, and the section on rolling back. Note the emphasis on rehearsing deployments and on rollback as a planned, tested operation. - The Site Reliability Workbook, Ch. 16, "Canarying Releases" (Free to read online)
What a canary is and is not, choosing canary population and duration, and evaluating canary metrics against a control group. - Site Reliability Engineering, Ch. 8, "Release Engineering" (Free to read online (CC BY-NC-ND 4.0))
Hermetic builds, the release process, and configuration management at Google. Compare it with the “Build once, and know what you built” notes. - Continuous Delivery, Ch. 12, "Managing Data" (Purchase; summaries are free on the book's site)
Database scripting and migrations, and decoupling application deployment from database migration.