Twelve-Factor processes are stateless and disposable. Anything that must survive belongs in a backing store. In containers, that store is a volume. The container's writable layer lives exactly as long as the container.
The lifecycle events differ:
| Event | Writable layer | Anonymous volume | Named volume |
|---|---|---|---|
restart |
kept | kept | kept |
recreate (up --force-recreate, new image) |
lost | usually reattached by Compose | kept |
down, then up |
lost | new, empty | kept |
down -v |
lost | deleted | deleted |
A durability test that only restarts proves almost nothing. Test the event your releases actually perform.
Surviving is half the requirement. The other half is how many times each item is delivered. A queue that persists its log but not its consumer offset redelivers after a crash, which is at-least-once delivery. One that advances the offset before the work is durable can drop items. "Exactly once" in practice means persisting progress atomically alongside, or after, the effect, plus consumers that are idempotent when they do see a duplicate. DDIA's stream processing chapter walks through these guarantees.
The SRE book's data integrity chapter adds the operator's rule: backups are worthless until a restore has been tested. The same holds for volumes. Prove recovery with the operation you fear.
A backup only counts once you have restored from it, and the restore is a decision too. A logical dump such as pg_dump holds the database as it was at one moment. Restoring it over the live database brings back what was lost and also throws away every write since, which is usually a second incident. The safer pattern restores into a separate database, compares it with the live one, and copies back exactly the rows that are missing. Keep the original ids, because other tables and other systems point at them. Before copying anything, read what the destructive change was meant to do, so you do not undo the part that was correct.