Skip to main content
RunBook Academy

Course catalogue

Every course, one place.

Each RunBook Academy course teaches the technology and how to operate it safely in production: how it fails, how to troubleshoot it, how to back it up, how to recover it, how to observe it, and how to secure it.

Suggested learning path

A newcomer who follows the cross-course recommendations in this order gets the dependencies right: every later course can be read with the production Linux background an earlier one assumed.

  1. 1Proxmox VE
  2. 2Observability

operating-system · security · storage

Linux for Production Sysadmins

Operate Linux on production hardware and Linux server clusters safely: shell, storage, networking, security, performance, observability, backup, DR, high availability, and incident response.

Lessons
523
Labs
68
Runbooks
26
Break/Fix
21
Lists
13

containers · automation · security

Docker & Containers for Production Sysadmins

Operate Docker Engine in production: from namespaces and cgroups through security hardening, observability, backups, incident response, and a complete capstone environment.

Lessons
237
Labs
27
Runbooks
24
Break/Fix
20
Lists
13

virtualisation · storage · networking · security

Proxmox VE for Production Operators

Operate Proxmox VE in production: planning, installation, networking, ZFS, Ceph, clustering, HA, PBS, security, monitoring, performance, maintenance, troubleshooting, CLI & automation, migration, and ops practice.

Lessons
253
Labs
23
Runbooks
24
Break/Fix
24
Lists
12

automation · security · operating-system

Ansible for Production Sysadmins

Design, secure, test and operate Ansible automation against production Linux fleets: inventories, idempotent roles, secrets, blast-radius control, canaries, rolling change, patching, and recovery when a run stops halfway.

Lessons
374
Labs
28
Runbooks
23
Break/Fix
22
Lists
13

automation · cloud · security

Terraform for Production Sysadmins

Safely design, implement, review, operate, troubleshoot and recover production Terraform infrastructure across business-critical environments.

Lessons
218
Labs
21
Runbooks
22
Break/Fix
26
Lists
12
Coming soon

observability · monitoring

Observability for Production Sysadmins

Design, deploy, secure, operate, scale, troubleshoot, and recover a production observability platform based on Prometheus, Grafana, Loki, and Tempo — and use it to operate the rest of production.

Lessons
684
Labs
30
Runbooks
30
Break/Fix
32
Lists
16