Skip to main content
RunBook Academy

Backup & DR · Curriculum

Curriculum

127 lessons across 24 parts. Lessons build on each other; later parts assume familiarity with earlier material.

Part I

Recovery Objectives, Vocabulary and Failure Models

Why recovery rather than backup is the objective, the seven words that are not synonyms, naming the failures you are actually protecting against, classifying data by whether it can be reconstructed, and building the first inventory and dependency map.

6 lessons
  1. 01Backup is not the objective — recovery isFoundations · foundation · ~26 minLab
  2. 02Seven words that are not synonymsFoundations · foundation · ~28 min
  3. 03The failure model: naming what you protect againstFoundations · foundation · ~27 min
  4. 04Classifying data by what it costs to loseFoundations · foundation · ~26 minLab
  5. 05What a backup system owes you, and what it never promisedFoundations · foundation · ~26 min
  6. 06Reading an estate: inventory and the first dependency mapFoundations · intermediate · ~30 min

Part II

RPO, RTO and Recovery Sequencing

Engineering an acceptable loss window, counting everything that happens between the incident and a working service, estimating restore duration from evidence, service tiers, the recovery dependency graph, and measuring the promise you made.

6 lessons
  1. 01RPO: engineering an acceptable loss windowObjectives · intermediate · ~30 minLab
  2. 02RTO: everything between the incident and a working serviceObjectives · intermediate · ~30 minLab
  3. 03Estimating restore duration from evidenceObjectives · intermediate · ~29 min
  4. 04Service tiers and what a tier actually buysObjectives · intermediate · ~26 min
  5. 05The recovery dependency graphObjectives · intermediate · ~30 min
  6. 06Recovery SLOs: measuring the promise you madeObjectives · intermediate · ~27 min

Part III

Backup Architecture: Copies, Chains, Retention and Capacity

3-2-1 read critically, independence of media and failure and security domains, full and differential and incremental and incremental-forever, restore dependency chains, retention as a pattern rather than a prescription, deduplication and compression, and capacity planning.

7 lessons
  1. 013-2-1 read criticallyArchitecture · intermediate · ~28 minLab
  2. 02Independence: media, failure domain, security domain, geographyArchitecture · intermediate · ~29 min
  3. 03Full, differential, incremental and incremental-foreverArchitecture · intermediate · ~28 minLab
  4. 04Restore dependency chains and the risk of chain lengthArchitecture · intermediate · ~28 min
  5. 05Retention schemes as patterns, not prescriptionsArchitecture · intermediate · ~27 minLab
  6. 06Deduplication and compression: the benefit and the billArchitecture · intermediate · ~28 minLab
  7. 07Capacity planning for a recovery estateArchitecture · intermediate · ~29 min

Part IV

Consistency, Integrity and Proof of Restorability

Crash-consistent versus application-consistent and the gap between them, open files and buffered writes and quiescing, checksums and what a completed job does not mean, repository verification, detecting corruption before a restore needs the data, and validation with teeth.

6 lessons
  1. 01Crash-consistent, application-consistent, and the gap between themConsistency · intermediate · ~30 minLab
  2. 02Filesystem consistency: open files, buffered writes and quiescingConsistency · intermediate · ~28 min
  3. 03Checksums, verification, and what a completed job does not meanConsistency · intermediate · ~27 min
  4. 04Repository verification and the cost of proving data intactConsistency · intermediate · ~28 minLab
  5. 05Detecting corruption before a restore needs the dataConsistency · advanced · ~28 min
  6. 06Proving a restore: validation with teethConsistency · advanced · ~29 minLab

Part V

Linux File-Level Backup and Restore

What actually has to be backed up on a Linux host, rsync and what it is not, mirrors that propagate deletion and corruption and ransomware, archive semantics for permissions and ACLs and extended attributes and sparse files, and restoring both a single file and a complete service.

6 lessons
  1. 01What actually has to be backed up on a Linux hostFiles · intermediate · ~28 minLab
  2. 02rsync: what it is, and what it is notFiles · intermediate · ~28 minLab
  3. 03Mirror propagation: deletion, corruption and ransomwareFiles · intermediate · ~27 min
  4. 04Archive semantics: permissions, ACLs, xattrs, sparse files and hard linksFiles · advanced · ~29 min
  5. 05Restoring a single file under pressureFiles · intermediate · ~26 minLab
  6. 06Restoring a complete service onto a clean hostFiles · advanced · ~30 min

Part VI

Snapshots: LVM, Btrfs and ZFS

Copy-on-write as the mechanism under every snapshot, LVM snapshot sizing and growth and invalidation, Btrfs subvolumes, ZFS snapshots and clones and space accounting, replication that turns a snapshot into a copy, the snapshot test, and rollback as an operation that destroys newer data.

7 lessons
  1. 01Copy-on-write: the mechanism under every snapshotSnapshots · intermediate · ~28 minLab
  2. 02LVM snapshots: sizing, growth and invalidationSnapshots · advanced · ~29 min
  3. 03Btrfs subvolumes and snapshotsSnapshots · advanced · ~28 min
  4. 04ZFS snapshots, clones and space accountingSnapshots · advanced · ~29 min
  5. 05Replication with send/receive: when a snapshot becomes a copySnapshots · advanced · ~29 minLab
  6. 06The snapshot test: why hourly snapshots may protect nothingSnapshots · advanced · ~28 min
  7. 07Rollback: the operation that destroys newer dataSnapshots · advanced · ~27 min

Part VII

Block Images, Bare-Metal Recovery and Reconstruction

Block-level images and their limits, bare-metal recovery as an architecture, the boot and partitioning state nobody backs up, the choice between rebuilding and restoring, what configuration management can and cannot reconstruct, and the drift that opens the reconstruction gap.

6 lessons
  1. 01Block-level images: use cases and limitsImages · intermediate · ~27 minLab
  2. 02Bare-metal recovery as an architectureImages · advanced · ~29 min
  3. 03Boot, partitioning and the parts nobody backs upImages · advanced · ~27 min
  4. 04Rebuild versus restore: the central choiceImages · advanced · ~28 min
  5. 05What configuration management can and cannot reconstructImages · advanced · ~28 min
  6. 06Golden images, drift and the reconstruction gapImages · intermediate · ~27 min

Part VIII

Backup Repositories: restic, Borg and Repository Failure

The chunk and index and snapshot model, restic and BorgBackup as two different sets of trade-offs, what each level of repository verification actually proves, corruption that survives a shallow check, losing the catalogue, and choosing a tool honestly.

7 lessons
  1. 01The repository model: chunks, indexes and snapshotsRepositories · intermediate · ~28 minLab
  2. 02restic in production: backup, retention and pruneRepositories · advanced · ~29 min
  3. 03BorgBackup: a different set of trade-offsRepositories · advanced · ~28 min
  4. 04Repository integrity: what each level of check provesRepositories · advanced · ~28 min
  5. 05Repository corruption and the partial restoreRepositories · advanced · ~29 min
  6. 06Losing the catalogue: metadata recoveryRepositories · advanced · ~28 minLab
  7. 07Choosing a backup tool honestlyRepositories · intermediate · ~27 min

Part IX

Encryption, Keys and Key Recovery

Encryption in transit and at rest for backup data, the paradox of a key that died with the environment it protected, key custody and escrow and split knowledge, KMS and HSM dependencies inside a recovery path, rotating without orphaning history, and proving you can still decrypt.

6 lessons
  1. 01Encryption in transit and at rest for backup dataEncryption · intermediate · ~27 min
  2. 02The encryption paradox: the key that died with productionEncryption · advanced · ~28 minLab
  3. 03Key custody, escrow and split knowledgeEncryption · advanced · ~29 min
  4. 04KMS and HSM dependencies inside a recovery pathEncryption · advanced · ~28 min
  5. 05Rotating backup encryption without orphaning historyEncryption · advanced · ~27 min
  6. 06Proving you can still decryptEncryption · advanced · ~27 min

Part X

Object Storage, Versioning and Retention

Why durability is not backup, versioning and delete markers and what a deletion actually did, lifecycle rules that quietly remove recovery points, governance and compliance retention modes, the delete permission as the control that matters, and what varies between S3-compatible implementations.

6 lessons
  1. 01Object storage as a backup target: durability is not backupObject storage · intermediate · ~27 minLab
  2. 02Versioning and delete markers: what a deletion actually didObject storage · advanced · ~28 min
  3. 03Lifecycle rules and the backups they quietly removeObject storage · advanced · ~27 min
  4. 04Retention modes: governance, compliance, and what each promisesObject storage · advanced · ~29 minLab
  5. 05Credentials, least privilege and the delete permissionObject storage · advanced · ~28 minLab
  6. 06What varies between S3-compatible implementationsObject storage · advanced · ~27 min

Part XI

Immutability, Air Gap and Ransomware Resilience

Whether an attacker who owns production can destroy the backups, object lock in practice, logical versus physical air gaps, separation of duties for backup administration, establishing a compromise timeline, selecting a clean recovery point, avoiding the restoration of persistence, and preserving evidence.

8 lessons
  1. 01Can the attacker delete the backups?Immutability · advanced · ~28 min
  2. 02Object lock in practice: what it stops, and what it does notImmutability · advanced · ~29 min
  3. 03Logical air gap and physical air gapImmutability · advanced · ~27 min
  4. 04Separation of duties and the backup administratorImmutability · advanced · ~28 min
  5. 05Ransomware: determining scope and the compromise timelineImmutability · advanced · ~30 minLab
  6. 06Selecting a clean recovery pointImmutability · advanced · ~29 min
  7. 07Recovering without restoring the persistence an attacker leftImmutability · advanced · ~29 minLab
  8. 08Preserving evidence while restoring serviceImmutability · advanced · ~27 min

Part XII

Virtual Machine and Hypervisor Recovery

VM snapshot against VM backup, guest quiescing and application consistency inside a guest, the estate state that is not a disk image, restoring a VM into a working network, hypervisor loss with the repository intact, and losing an entire virtualisation cluster.

6 lessons
  1. 01VM snapshot, VM backup, and the difference that mattersVirtualisation · intermediate · ~27 min
  2. 02Guest quiescing and application consistency inside a VMVirtualisation · advanced · ~28 min
  3. 03Backing up a virtualisation estate: what else is stateVirtualisation · advanced · ~27 min
  4. 04Restoring a VM: the parts that are not the disk imageVirtualisation · advanced · ~28 min
  5. 05Hypervisor loss with the repository intactVirtualisation · advanced · ~28 min
  6. 06Losing an entire virtualisation clusterVirtualisation · advanced · ~29 minLab

Part XIII

Container and Kubernetes Recovery

Why an image is not a backup, where container state actually lives, backing up volumes with consistency, the four separate things a Kubernetes backup has to cover, etcd snapshot and restore, persistent volumes and CSI snapshots, the boundary of Velero, and recovery onto a clean cluster.

8 lessons
  1. 01A container image is not a backupContainers · intermediate · ~27 minLab
  2. 02Volumes, bind mounts, and where the state actually isContainers · intermediate · ~27 min
  3. 03Backing up container data with consistencyContainers · advanced · ~28 min
  4. 04The Kubernetes backup model: four separate thingsKubernetes · advanced · ~29 minLab
  5. 05etcd: snapshot and restoreKubernetes · advanced · ~30 minLab
  6. 06Persistent volumes and CSI snapshotsKubernetes · advanced · ~28 min
  7. 07Velero: what it covers and what it does notKubernetes · advanced · ~27 min
  8. 08Recovering a cluster onto clean infrastructureKubernetes · advanced · ~30 min

Part XIV

Database Backup and Point-in-Time Recovery

Why copying live database files is not a backup even when it appears to work, logical against physical backups, WAL archiving and the continuity of the archive, performing a point-in-time recovery to a chosen target, recovering from a logical mistake, and validating the result.

6 lessons
  1. 01Why copying live database files is not a backupDatabases · advanced · ~29 min
  2. 02Logical and physical backups comparedDatabases · advanced · ~28 min
  3. 03WAL archiving and the continuity of the archiveDatabases · advanced · ~29 minLab
  4. 04Performing a point-in-time recoveryDatabases · advanced · ~30 min
  5. 05Recovering from a logical mistakeDatabases · advanced · ~28 min
  6. 06Validating a recovered databaseDatabases · advanced · ~28 minLab

Part XV

Infrastructure Reconstruction: IaC, Config, Network and Identity

Code rebuilds infrastructure while backup restores state, Terraform state as critical recovery material, the boundary of configuration management, Git and pipelines and artifacts, network device and firewall configuration, DNS and DHCP and IPAM, and the identity and PKI bootstrap problem.

8 lessons
  1. 01Code rebuilds infrastructure; backup restores stateReconstruction · intermediate · ~45 minLab
  2. 02Terraform state as critical recovery materialReconstruction · advanced · ~50 min
  3. 03Ansible and the boundary of reconstructionReconstruction · intermediate · ~45 min
  4. 04Git, pipelines and artifacts in a recoveryReconstruction · intermediate · ~45 min
  5. 05Network device configuration recoveryNetwork and identity · intermediate · ~45 min
  6. 06Firewall and routing recoveryNetwork and identity · advanced · ~50 min
  7. 07DNS, DHCP and IPAM as recovery dependenciesNetwork and identity · intermediate · ~45 min
  8. 08Identity, secrets and PKI: the bootstrap problemNetwork and identity · expert · ~55 minLab

Part XVI

Monitoring, Restore Testing and Recovery Assurance

The signals that indicate recovery capability, restore-point age instead of job success, alerts worth waking someone for, automated restore verification, a testing maturity progression, restoring onto clean infrastructure to expose hidden dependencies, business-level validation, and treating a failed restore as an incident.

8 lessons
  1. 01Monitoring the backup estate: the signals that matterSignals · intermediate · ~45 minLab
  2. 02Restore-point age, not job successSignals · intermediate · ~45 min
  3. 03Alerting on recovery capabilitySignals · advanced · ~50 min
  4. 04Automated restore verificationVerification · advanced · ~55 minLab
  5. 05Recovery testing: a maturity progressionVerification · advanced · ~45 minLab
  6. 06Restoring onto clean infrastructure to expose hidden dependenciesVerification · advanced · ~50 min
  7. 07Business-level validationVerification · advanced · ~45 min
  8. 08A failed restore is a production incidentVerification · intermediate · ~45 minLab

Part XVII

Disaster Recovery: Failover, Failback and the Recovery Estate

Why disaster recovery is not high availability, what cold and warm and hot actually cost, active/passive against active/active, DNS and routing and certificates during a failover, secrets and identity and third-party dependencies, executing a failover, the data created in the recovery site, failback as a second outage, exercises, and human factors.

10 lessons
  1. 01Disaster recovery is not high availabilityFoundations · intermediate · ~45 min
  2. 02Cold, warm and hot: what each actually costsFoundations · intermediate · ~45 min
  3. 03Active/passive and active/active for disaster recoveryFoundations · advanced · ~50 min
  4. 04DNS, routing and certificates during a failoverExecution · advanced · ~50 minLab
  5. 05Secrets, identity and third-party dependencies in DRExecution · expert · ~55 min
  6. 06Executing a failoverExecution · advanced · ~50 minLab
  7. 07Operating in DR, and the data you create thereExecution · advanced · ~50 minLab
  8. 08Failback: the second outageExecution · expert · ~55 min
  9. 09DR exercises from tabletop to fullAssurance · intermediate · ~45 minLab
  10. 10Incident command and human factors during recoveryAssurance · intermediate · ~45 min

Part XVIII

Backup Platform DR, Media, Cost and Compliance

Recovering when the backup platform is what failed, rebuilding a catalogue, backup network design and restore throughput, what tape is still good at, cloud archive tiers whose retrieval breaks an RTO, media lifecycle and decommissioning, cost engineering, retention and sovereignty, and a defensible reference architecture.

9 lessons
  1. 01When the backup platform is the thing that failedPlatform recovery · advanced · ~50 min
  2. 02Rebuilding a backup cataloguePlatform recovery · advanced · ~50 min
  3. 03Backup network design and restore throughputMedia and throughput · advanced · ~50 minLab
  4. 04Tape: what it is still good atMedia and throughput · intermediate · ~45 minLab
  5. 05Cloud archive tiers and the retrieval that breaks RTOMedia and throughput · advanced · ~50 minLab
  6. 06Media lifecycle, bit rot and decommissioningMedia and throughput · intermediate · ~45 min
  7. 07Cost engineering for a recovery estateGovernance · advanced · ~50 min
  8. 08Retention, legal hold and data sovereigntyGovernance · advanced · ~50 min
  9. 09A defensible reference architectureGovernance · expert · ~55 minLab

Part Labs

Labs

Disposable-environment labs covering classification and RPO/RTO calculation, file restore, rsync propagation, archive semantics, LVM and ZFS and Btrfs snapshots, repository integrity and corruption, object storage versioning and object lock, container and Kubernetes and etcd recovery, database point-in-time recovery, key recovery, automated restore verification, a measured RTO, and a failover with failback.

0 lessons

No lessons published in this part yet. The full curriculum is planned in docs/courses/backup-dr/curriculum.md on GitHub.

Part Runbooks

Runbooks

Operational procedures for restoring files, services, virtual machines, container and Kubernetes data and databases, for investigating failed backups and failed restores and repository corruption and capacity exhaustion, for recovering encryption keys and the backup control plane, and for ransomware recovery, failover, validation, failback and total site loss.

0 lessons

Part Checklists

Checklists

Production-readiness reviews for backup and restore readiness, RPO and RTO, ransomware resilience, immutability, repository security, database and virtual machine and Kubernetes and network device recovery, secrets and PKI recovery, pre-restore and post-restore validation, and disaster-recovery site and exercise readiness.

0 lessons

Part Breakfix

Break/Fix Scenarios

Evidence-first diagnosis of green jobs that cannot restore, exhausted repositories, unavailable keys, broken incremental chains, corruption that passes a shallow check, lifecycle rules that deleted the recovery points, inconsistent backups, unreachable recovery targets, and the dependency failures that stop a technically successful recovery from restoring the service.

0 lessons

Part Capstone

Production Capstone

Design, build and defend a complete recovery estate for a realistic production site: classified data, engineered RPO and RTO, immutable and offsite copies, recoverable keys, proven restores, monitored recovery capability, a rehearsed failover and failback, and twelve injected incidents to survive.

1 lesson
  1. 01Capstone — design, build and defend a production recovery estateCapstone · expert · ~300 min

Part Final

Final Assessment

Final theory assessment of operational judgement, plus a final practical assessment of an inherited backup and disaster-recovery estate carrying realistic coverage, consistency, immutability, key-custody, retention and recovery-time defects.

0 lessons

No lessons published in this part yet. The full curriculum is planned in docs/courses/backup-dr/curriculum.md on GitHub.