Proxmox VE · Operational reference
Toolkit
Operational artefacts for Proxmox VE: checklists for cadences, runbooks for incidents, hands-on labs for skills, and break/fix scenarios for troubleshooting reflexes. None of this replaces reading the corresponding lessons — these are the artefacts you keep open in a second tab during real work.
- Runbooks
- 24
- Checklists
- 12
- Labs
- 23
- Break/Fix
- 24
- Assessments
- 1
Operational checklists
Repetitive tasks performed on a schedule.
Before deployment
Weekly
Quarterly
Runbooks
Step-by-step procedures. Treat the rollback as part of the procedure — never skip it.
Critical risk
- criticalcluster affectingRecover a Ceph cluster from a full or near-full OSD~180 min · 11 steps · verified 2026-08-12→
- criticalcluster affectingRecover a cluster from a corrupt /etc/pve (pmxcfs)~120 min · 12 steps · verified 2026-08-12→
- criticalcluster affectingUpgrade a cluster from PVE 8 to PVE 9, node by node~480 min · 14 steps · verified 2026-08-12→
High risk
- highdata loss riskExpand a ZFS pool safely~120 min · 12 steps · verified 2026-08-12→
- highcluster affectingRecover the SDN configuration after a bad apply~120 min · 11 steps · verified 2026-08-12→
- highcluster affectingRecover quorum after losing a single node~30 min · 7 steps · verified 2026-08-07→
- highcluster affectingPermanently remove a node from a cluster without breaking quorum~90 min · 14 steps · verified 2026-08-12→
- highcluster affectingReplace a failed OSD in Ceph~60 min · 5 steps · verified 2026-08-07→
- highcluster affectingReplace a failed node with a rebuilt one of the same name~240 min · 13 steps · verified 2026-08-12→
- highcluster affectingRespond to a fencing event and a self-fenced node~90 min · 11 steps · verified 2026-08-12→
- highdata loss riskRestore a corrupted ZFS dataset from PBS~90 min · 8 steps · verified 2026-08-07→
Medium risk
- mediumcluster affectingAdd a new PVE node to an existing cluster~30 min · 5 steps · verified 2026-08-07→
- mediumservice affectingDiagnose and fix memory pressure on a PVE host~30 min · 6 steps · verified 2026-08-07→
- mediumservice affectingEvacuate a node for planned maintenance~60 min · 9 steps · verified 2026-08-07→
- mediumcluster affectingExpand Ceph capacity by adding OSDs, without a rebalance storm~180 min · 13 steps · verified 2026-08-12→
- mediumservice affectingLive-migrate a running VM, and recover when it stalls~45 min · 10 steps · verified 2026-08-12→
- mediumdata loss riskMove a VM disk to different storage while it runs, and roll back a stalled move~90 min · 9 steps · verified 2026-08-12→
- mediumcluster affectingPut a node into maintenance~60 min · 9 steps · verified 2026-08-07→
- mediumservice affectingReplace a failed disk in a ZFS mirror~60 min · 5 steps · verified 2026-08-07→
- mediumservice affectingRestore a VM from PBS backup~30 min · 4 steps · verified 2026-08-07→
- mediumservice affectingRotate cluster and API credentials, including tokens~120 min · 11 steps · verified 2026-08-12→
- mediumservice affectingProve a PBS backup is genuinely restorable, not merely green~150 min · 13 steps · verified 2026-08-12→
Hands-on labs
Time-boxed exercises.
A · Physical hardware
- Configure bond + VLAN + bridge for VM traffic~45 min · 3 objectives→
- Set up a 3-node Ceph cluster and create a CephFS share~90 min · 4 objectives→
- Build a 3-node Ceph cluster and create OSD-backed storage~180 min · 4 objectives→
- Build a hardened LXC container with cloud-init and SSH keys~45 min · 4 objectives→
- Provision a new cluster member and migrate workloads to it~60 min · 4 objectives→
- Provision a 3-node cluster from bare metal~120 min · 4 objectives→
- Configure email alerts for HA events and Ceph health~45 min · 4 objectives→
- Configure PVE firewall with zone-based isolation between VMs~60 min · 4 objectives→
- Configure HA for a VM and verify failover~60 min · 4 objectives→
- Enable HA on a VM and run a controlled failover test~60 min · 4 objectives→
- Configure iSCSI target and connect via multipath on PVE~90 min · 4 objectives→
- Upgrade a PVE node in place from 9.2 to 9.3 with no downtime~90 min · 4 objectives→
- Set up Prometheus, node-exporter, and Grafana for cluster monitoring~90 min · 4 objectives→
- Configure NFS storage for VM workloads and live migration~45 min · 4 objectives→
- Install PBS, configure a datastore, and run a first backup~60 min · 4 objectives→
- Deploy PBS, configure datastore, run first backup and restore test~75 min · 4 objectives→
- Pass through a PCIe device and verify it appears inside the guest~60 min · 4 objectives→
- Set up SDN with a VLAN zone and verify isolation~30 min · 4 objectives→
- Configure SDN with two zones and verify routing between them~75 min · 4 objectives→
- Enable two-factor authentication for a PVE user~45 min · 4 objectives→
- Right-size a VM using perf, ballooning, and CPU pinning~75 min · 4 objectives→
- Set up ZFS mirror with spare and exercise disk replacement~60 min · 4 objectives→
- Configure ZFS replication and run a planned failover~75 min · 4 objectives→
Break/Fix scenarios
Troubleshooting drills. Each scenario gives symptoms and evidence, then hides the solution behind a reveal.
performance
ceph
security
corosync
network
- beginnerDNS resolution from VMs fails after enabling Pi-hole~10 min · 4 symptoms→
- beginnerNetwork becomes unreachable after enabling the PVE firewall~10 min · 4 symptoms→
- intermediateNFS share becomes read-only during heavy writes~15 min · 4 symptoms→
- beginnerNetwork between PVE and PBS becomes unreachable~15 min · 4 symptoms→
storage
pbs
Assessments
Production-readiness examinations. Each combines auto-scored questions with scenario-based rubrics you can use to grade yourself.