Skip to main content
RunBook Academy

Course catalogue

Every course, one place.

Each RunBook Academy course teaches the technology and how to operate it safely in production: how it fails, how to troubleshoot it, how to back it up, how to recover it, how to observe it, and how to secure it.

Suggested learning path

A newcomer who follows the cross-course recommendations in this order gets the dependencies right: every later course can be read with the production Linux background an earlier one assumed.

  1. 1OPNsense
  2. 2Backup & DR

operating-system · security · storage

Linux for Production Sysadmins

Operate Linux on production hardware and Linux server clusters safely: shell, storage, networking, security, performance, observability, backup, DR, high availability, and incident response.

Lessons
523
Labs
68
Runbooks
26
Break/Fix
21
Lists
13

containers · automation · security

Docker & Containers for Production Sysadmins

Operate Docker Engine in production: from namespaces and cgroups through security hardening, observability, backups, incident response, and a complete capstone environment.

Lessons
237
Labs
27
Runbooks
24
Break/Fix
20
Lists
13

virtualisation · storage · networking · security

Proxmox VE for Production Operators

Operate Proxmox VE in production: planning, installation, networking, ZFS, Ceph, clustering, HA, PBS, security, monitoring, performance, maintenance, troubleshooting, CLI & automation, migration, and ops practice.

Lessons
253
Labs
23
Runbooks
24
Break/Fix
24
Lists
12

automation · security · operating-system

Ansible for Production Sysadmins

Design, secure, test and operate Ansible automation against production Linux fleets: inventories, idempotent roles, secrets, blast-radius control, canaries, rolling change, patching, and recovery when a run stops halfway.

Lessons
374
Labs
28
Runbooks
23
Break/Fix
22
Lists
13

automation · cloud · security

Terraform for Production Sysadmins

Safely design, implement, review, operate, troubleshoot and recover production Terraform infrastructure across business-critical environments.

Lessons
179
Labs
21
Runbooks
22
Break/Fix
26
Lists
12

Design, deploy, secure, operate, scale, troubleshoot, and recover a production observability platform based on Prometheus, Grafana, Loki, and Tempo — and use it to operate the rest of production.

Lessons
684
Labs
30
Runbooks
30
Break/Fix
32
Lists
16

networking · security · operating-system

VyOS for Production Network Engineers

Design, deploy, secure, automate, operate, troubleshoot and recover VyOS routers and routing clusters supporting mission-critical production networks.

Lessons
342
Labs
25
Runbooks
39
Break/Fix
43
Lists
18

containers · automation · cloud · networking · storage

Kubernetes for Production Sysadmins

Design, deploy, secure, operate, monitor, troubleshoot, upgrade and recover Kubernetes clusters supporting business-critical production workloads.

Lessons
787
Labs
25
Runbooks
25
Break/Fix
30
Lists
15

Design, deploy, secure, operate, monitor, troubleshoot, scale, upgrade, and recover production Ceph clusters: RADOS, CRUSH, MON/MGR/OSD, RBD, CephFS, RGW, capacity, performance, recovery, Proxmox and Kubernetes integration.

Lessons
744
Labs
31
Runbooks
40
Break/Fix
41
Lists
18

Use Git safely, design production CI/CD pipelines, secure software and infrastructure supply chains, and operate GitOps workflows for business-critical infrastructure.

Lessons
716
Labs
31
Runbooks
30
Break/Fix
30
Lists
15
Beta / reference release

storage · operating-system · automation

PostgreSQL for Production Sysadmins

Operate PostgreSQL where availability, durability and recovery matter: architecture, MVCC and vacuum, locks, WAL, backup and point-in-time recovery, replication, failover and disaster recovery.

Lessons
127
Labs
26
Runbooks
26
Break/Fix
30
Lists
15