wtf( )unctionsystem design, drawn
← all problemsCloud PracticeEasy

Backups every night, restores never

The nightly backup has succeeded every night for over a year, and the dashboard is green.

When a restore was finally needed it did not work. The dump excluded a schema added eight months ago, the encryption key had been rotated since, and nobody knew how long a restore takes — so the recovery objective everyone had agreed to was a number with nothing behind it.

Add the step that turns a backup into a recovery capability.
Components — tap one, then tap a slot on the diagram
!The backup job succeeded 400 nights running. The restore failed, because nobody had ever run one.

Outside every boundary: Production (the live data), RTO: 4 hours (never measured; FAILED: unverified), Backup storage (a year of dumps), Nightly backup (400 successes), an empty slot for the on a schedule, automatically Connections: Production calls Nightly backup (step 1) Nightly backup calls Backup storage (step 2) Backup storage calls on a schedule, automatically (step 3) on a schedule, automatically controls RTO: 4 hours — measured, not assumed (step 4)

Productionthe live data
RTO: 4 hoursnever measuredunverified
Backup storagea year of dumps
Nightly backup400 successes