Backup & DR · Operational reference
Toolkit
Operational artefacts for Backup & DR: checklists for cadences, runbooks for incidents, hands-on labs for skills, and break/fix scenarios for troubleshooting reflexes. None of this replaces reading the corresponding lessons — these are the artefacts you keep open in a second tab during real work.
- Runbooks
- 26
- Checklists
- 15
- Labs
- 26
- Break/Fix
- 30
- Assessments
- 2
Operational checklists
Repetitive tasks performed on a schedule.
Quarterly
- Backup production readiness20 items→
- Restore readiness21 items→
- Ransomware resilience21 items→
- Immutable backup review20 items→
- Backup repository security20 items→
- Database backup readiness20 items→
- Virtual machine backup readiness21 items→
- Kubernetes recovery readiness20 items→
- Network device recovery readiness21 items→
- Secrets and PKI recovery readiness20 items→
- DR site readiness22 items→
Runbooks
Step-by-step procedures. Treat the rollback as part of the procedure — never skip it.
Critical risk
- criticalcluster affectingRecover Kubernetes cluster state~120 min · 13 steps · verified 2026-08-28→
- criticaldata loss riskExecute a PostgreSQL point-in-time recovery~150 min · 14 steps · verified 2026-08-28→
- criticalservice affectingRestore network device configuration~60 min · 15 steps · verified 2026-08-28→
- criticalcluster affectingRecover the backup control plane~150 min · 17 steps · verified 2026-08-28→
- criticalsecurity relevantRecover encryption key access~90 min · 16 steps · verified 2026-08-28→
- criticalsecurity relevantExecute ransomware recovery~240 min · 17 steps · verified 2026-08-28→
- criticalsecurity relevantSelect a clean recovery point~90 min · 16 steps · verified 2026-08-28→
- criticalservice affectingExecute a DR failover~180 min · 17 steps · verified 2026-08-28→
- criticaldata loss riskExecute a DR failback~180 min · 21 steps · verified 2026-08-28→
- criticalcluster affectingRecover after complete primary-site loss~300 min · 17 steps · verified 2026-08-28→
High risk
- highservice affectingRestore a virtual machine~90 min · 15 steps · verified 2026-08-28→
- highservice affectingRecover Kubernetes application data~90 min · 14 steps · verified 2026-08-28→
- highdata loss riskRestore a PostgreSQL logical backup~75 min · 15 steps · verified 2026-08-28→
- highdata loss riskRecover an accidentally deleted dataset~75 min · 15 steps · verified 2026-08-28→
- highdata loss riskRecover a backup repository~120 min · 17 steps · verified 2026-08-28→
- highdata loss riskInvestigate a failed restore~60 min · 16 steps · verified 2026-08-28→
- highdata loss riskInvestigate repository corruption~75 min · 16 steps · verified 2026-08-28→
- highcluster affectingRebuild infrastructure from IaC plus restored state~150 min · 18 steps · verified 2026-08-28→
Medium risk
- mediumservice affectingRestore a deleted file~30 min · 13 steps · verified 2026-08-28→
- mediumservice affectingRestore a complete Linux service~90 min · 16 steps · verified 2026-08-28→
- mediumservice affectingRestore container persistent data~60 min · 13 steps · verified 2026-08-28→
- mediumdata loss riskInvestigate a failed backup job~45 min · 17 steps · verified 2026-08-28→
- mediumdata loss riskInvestigate backup capacity exhaustion~50 min · 18 steps · verified 2026-08-28→
- mediumservice affectingRetrieve an offsite or immutable copy~90 min · 16 steps · verified 2026-08-28→
- mediumservice affectingValidate a recovered application~60 min · 12 steps · verified 2026-08-28→
Hands-on labs
Time-boxed exercises.
C · Simulation
B · Nested virtualisation
- Back up and restore Linux file data, and prove the bytes~55 min · 6 objectives→
- Demonstrate rsync deletion and ransomware propagation~50 min · 6 objectives→
- Archive fidelity: what default flags silently discard~60 min · 5 objectives→
- LVM snapshots: copy-on-write growth, overflow and invalidation~70 min · 6 objectives→
- ZFS snapshots, clones and replication to an independent pool~70 min · 5 objectives→
- Btrfs subvolumes, read-only snapshots and send/receive~60 min · 6 objectives→
- Restore a deleted configuration file under time pressure~45 min · 6 objectives→
- Restore a complete service onto clean infrastructure~80 min · 6 objectives→
- Build a restic repository, restore it, and prove the bytes~60 min · 7 objectives→
- Detect repository corruption and measure the damage~65 min · 6 objectives→
- Borg append-only: what it stops, and recovering when it does not~75 min · 8 objectives→
- Write backups to object storage and restore from it~60 min · 7 objectives→
- Versioning and delete markers under a hostile identity~55 min · 7 objectives→
- Configure object lock and attempt deletion as production~65 min · 7 objectives→
- Governance versus compliance retention: measure the difference~50 min · 6 objectives→
- Back up and restore Docker volume data, and find what was missed~55 min · 6 objectives→
- Snapshot and restore etcd, and measure what the RPO cost~70 min · 7 objectives→
- Recover Kubernetes application data onto a rebuilt namespace~80 min · 7 objectives→
- PostgreSQL logical backup and restore, validated against an invariant~70 min · 7 objectives→
- PostgreSQL point-in-time recovery to a target before a mistake~100 min · 7 objectives→
- Prove the encryption key still opens the repository after losing production~75 min · 7 objectives→
- Automated restore verification that fails when the data is wrong~85 min · 7 objectives→
- Measure an actual RTO end to end, stage by stage~90 min · 7 objectives→
Break/Fix scenarios
Troubleshooting drills. Each scenario gives symptoms and evidence, then hides the solution behind a reveal.
bdr-restore-failure
bdr-capacity
bdr-encryption-key
bdr-credential
bdr-mirror-propagation
bdr-snapshot
bdr-repository-corruption
bdr-lifecycle
bdr-immutability
bdr-vm-consistency
bdr-db-consistency
bdr-wal-archive
bdr-pitr
bdr-etcd-restore
bdr-volume-snapshot
bdr-vm-restore
bdr-network-config
bdr-dr-secret
bdr-dr-dns
bdr-dr-firewall
bdr-rto-breach
bdr-archive-retrieval
bdr-replication-propagation
bdr-credential-separation
bdr-monitoring-blind
bdr-failback
bdr-business-validation
Assessments
Production-readiness examinations. Each combines auto-scored questions with scenario-based rubrics you can use to grade yourself.
- Backup, Restore & Disaster Recovery for Production Infrastructure — Final Practical Assessment~180 min · 12 graded questions · pass ≥ 80% · verified 2026-08-28→
- Backup, Restore & Disaster Recovery for Production Infrastructure — Final Theory Assessment~90 min · 25 graded questions · pass ≥ 80% · verified 2026-08-28→