Operate Linux on production hardware and Linux server clusters safely: shell, storage, networking, security, performance, observability, backup, DR, high availability, and incident response.
- Lessons
- 523
- Labs
- 68
- Runbooks
- 26
- Break/Fix
- 21
- Lists
- 13
Operate Docker Engine in production: from namespaces and cgroups through security hardening, observability, backups, incident response, and a complete capstone environment.
- Lessons
- 237
- Labs
- 27
- Runbooks
- 24
- Break/Fix
- 20
- Lists
- 13
Operate Proxmox VE in production: planning, installation, networking, ZFS, Ceph, clustering, HA, PBS, security, monitoring, performance, maintenance, troubleshooting, CLI & automation, migration, and ops practice.
- Lessons
- 253
- Labs
- 23
- Runbooks
- 24
- Break/Fix
- 24
- Lists
- 12
Design, secure, test and operate Ansible automation against production Linux fleets: inventories, idempotent roles, secrets, blast-radius control, canaries, rolling change, patching, and recovery when a run stops halfway.
- Lessons
- 374
- Labs
- 28
- Runbooks
- 23
- Break/Fix
- 22
- Lists
- 13
Safely design, implement, review, operate, troubleshoot and recover production Terraform infrastructure across business-critical environments.
- Lessons
- 179
- Labs
- 21
- Runbooks
- 22
- Break/Fix
- 26
- Lists
- 12
Design, deploy, secure, operate, scale, troubleshoot, and recover a production observability platform based on Prometheus, Grafana, Loki, and Tempo — and use it to operate the rest of production.
- Lessons
- 684
- Labs
- 30
- Runbooks
- 30
- Break/Fix
- 32
- Lists
- 16
Design, deploy, secure, operate, troubleshoot and recover OPNsense firewalls and OPNsense HA clusters supporting mission-critical production networks.
- Lessons
- 288
- Labs
- 21
- Runbooks
- 30
- Break/Fix
- 20
- Lists
- 15
Design, deploy, secure, automate, operate, troubleshoot and recover VyOS routers and routing clusters supporting mission-critical production networks.
- Lessons
- 342
- Labs
- 25
- Runbooks
- 39
- Break/Fix
- 43
- Lists
- 18
Design, deploy, secure, operate, monitor, troubleshoot, upgrade and recover Kubernetes clusters supporting business-critical production workloads.
- Lessons
- 787
- Labs
- 25
- Runbooks
- 25
- Break/Fix
- 30
- Lists
- 15
Design, deploy, secure, operate, monitor, troubleshoot, scale, upgrade, and recover production Ceph clusters: RADOS, CRUSH, MON/MGR/OSD, RBD, CephFS, RGW, capacity, performance, recovery, Proxmox and Kubernetes integration.
- Lessons
- 744
- Labs
- 31
- Runbooks
- 40
- Break/Fix
- 41
- Lists
- 18
Use Git safely, design production CI/CD pipelines, secure software and infrastructure supply chains, and operate GitOps workflows for business-critical infrastructure.
- Lessons
- 716
- Labs
- 31
- Runbooks
- 30
- Break/Fix
- 30
- Lists
- 15
Beta / reference releaseOperate credentials, keys, certificates and trust relationships safely: build and run a PKI, automate certificate issuance, manage secrets and short-lived credentials, and recover from compromise.
- Lessons
- 120
- Labs
- 26
- Runbooks
- 26
- Break/Fix
- 30
- Lists
- 15
Operate PostgreSQL where availability, durability and recovery matter: architecture, MVCC and vacuum, locks, WAL, backup and point-in-time recovery, replication, failover and disaster recovery.
- Lessons
- 127
- Labs
- 26
- Runbooks
- 26
- Break/Fix
- 30
- Lists
- 15
Backup is not the objective; recovery is. Design, secure, operate, test and execute the recovery of production infrastructure — from a deleted file to a lost datacentre.
- Lessons
- 127
- Labs
- 26
- Runbooks
- 26
- Break/Fix
- 30
- Lists
- 15