Ceph · Operational reference
Toolkit
Operational artefacts for Ceph: checklists for cadences, runbooks for incidents, hands-on labs for skills, and break/fix scenarios for troubleshooting reflexes. None of this replaces reading the corresponding lessons — these are the artefacts you keep open in a second tab during real work.
- Runbooks
- 40
- Checklists
- 18
- Labs
- 31
- Break/Fix
- 41
- Assessments
- 2
Operational checklists
Repetitive tasks performed on a schedule.
Before deployment
- Backup and DR Readiness Checklist8 items→
- CephFS Production Readiness Checklist6 items→
- Ceph Hardware Readiness Checklist6 items→
- Kubernetes + Ceph Readiness Checklist6 items→
- Ceph Network Readiness Checklist7 items→
- Ceph Observability Readiness Checklist6 items→
- Pool Design Checklist7 items→
- Pre-Upgrade Checklist10 items→
- Ceph Production Readiness Checklist9 items→
- Proxmox + Ceph Readiness Checklist6 items→
- RBD Production Readiness Checklist6 items→
- RGW Production Readiness Checklist6 items→
- Ceph Security Readiness Checklist8 items→
Runbooks
Step-by-step procedures. Treat the rollback as part of the procedure — never skip it.
Critical risk
- criticaldata loss riskExecute the Ceph disaster-recovery procedure~240 min · 8 steps · verified 2026-08-17→
- criticalcluster affectingRecover a full Ceph cluster~60 min · 5 steps · verified 2026-08-17→
- criticaldata loss riskRecover monitor quorum after multi-MON loss~90 min · 5 steps · verified 2026-08-17→
- criticalsecurity relevantRespond to Ceph credential compromise~30 min · 6 steps · verified 2026-08-17→
- criticaldata loss riskRoll back a failed upgrade~120 min · 6 steps · verified 2026-08-17→
High risk
- highcluster affectingInvestigate monitor quorum~20 min · 5 steps · verified 2026-08-17→
- highdata loss riskRestore an RBD workload~45 min · 7 steps · verified 2026-08-17→
- highdata loss riskRecover from a complete OSD host loss~120 min · 7 steps · verified 2026-08-17→
- highcluster affectingRecover a lost MON~60 min · 6 steps · verified 2026-08-17→
- highdata loss riskRecover from multiple OSD failures~90 min · 6 steps · verified 2026-08-17→
- highdata loss riskSafely remove a storage host~90 min · 6 steps · verified 2026-08-17→
- highdata loss riskRemove an OSD safely~60 min · 6 steps · verified 2026-08-17→
- highdata loss riskReplace a failed OSD end to end~90 min · 8 steps · verified 2026-08-17→
- highcluster affectingRestore Ceph configuration from backup~60 min · 6 steps · verified 2026-08-17→
- highdata loss riskPerform a Ceph upgrade safely~180 min · 8 steps · verified 2026-08-17→
Medium risk
- mediumservice affectingAdd Ceph capacity safely~45 min · 6 steps · verified 2026-08-17→
- mediumservice affectingAdd a new OSD host to a running cluster~45 min · 5 steps · verified 2026-08-17→
- mediumservice affectingAdd an OSD to an existing host~30 min · 5 steps · verified 2026-08-17→
- mediumsecurity relevantRotate a cephx key~45 min · 6 steps · verified 2026-08-17→
- mediumservice affectingDeploy a new Ceph cluster with cephadm~90 min · 8 steps · verified 2026-08-17→
- mediumdata loss riskInvestigate an inconsistent PG~30 min · 7 steps · verified 2026-08-17→
- mediumservice affectingInvestigate nearfull on a pool~20 min · 6 steps · verified 2026-08-17→
- mediumservice affectingInvestigate network latency on a Ceph cluster~30 min · 6 steps · verified 2026-08-17→
- mediumservice affectingInvestigate disk latency on a Ceph host~25 min · 7 steps · verified 2026-08-17→
- mediumservice affectingInvestigate slow ops across the cluster~25 min · 5 steps · verified 2026-08-17→
- mediumservice affectingPerform planned network maintenance safely~60 min · 6 steps · verified 2026-08-17→
- mediumservice affectingPerform planned node maintenance safely~120 min · 6 steps · verified 2026-08-17→
- mediumdata loss riskBack up an RBD workload~30 min · 6 steps · verified 2026-08-17→
- mediumservice affectingRecover a failed MDS daemon~30 min · 4 steps · verified 2026-08-17→
- mediumservice affectingRecover a failed MGR daemon~20 min · 5 steps · verified 2026-08-17→
- mediumservice affectingTroubleshoot Kubernetes / Ceph CSI~30 min · 6 steps · verified 2026-08-17→
- mediumservice affectingTroubleshoot Proxmox / Ceph storage performance~30 min · 5 steps · verified 2026-08-17→
- mediumservice affectingTune recovery and backfill safely~20 min · 8 steps · verified 2026-08-17→
Low risk
- lowcluster affectingBackup Ceph configuration and keyrings~20 min · 6 steps · verified 2026-08-17→
- lowservice affectingInvestigate a degraded PG~15 min · 6 steps · verified 2026-08-17→
- lowservice affectingInvestigate an OSD down event~20 min · 9 steps · verified 2026-08-17→
- lowservice affectingTroubleshoot a CephFS mount~20 min · 5 steps · verified 2026-08-17→
- lowservice affectingTroubleshoot an RBD client~15 min · 7 steps · verified 2026-08-17→
- lowservice affectingTroubleshoot a RGW client~15 min · 5 steps · verified 2026-08-17→
- lowservice affectingValidate the cluster after a major recovery~30 min · 5 steps · verified 2026-08-17→
Hands-on labs
Time-boxed exercises.
B · Nested virtualisation
- Lab 1: Deploy a Ceph cluster with cephadm~120 min · 3 objectives→
- Lab 2: OSD lifecycle~90 min · 4 objectives→
- Lab 3: RBD images~90 min · 5 objectives→
- Lab 4: CephFS~90 min · 4 objectives→
- Lab 5: RGW users and buckets~90 min · 4 objectives→
- Lab 6: CRUSH~90 min · 4 objectives→
- Lab 7: Replicated vs erasure-coded pools~90 min · 4 objectives→
- Lab 8: Ceph networking~75 min · 4 objectives→
- Lab 9: cephx capabilities~60 min · 4 objectives→
- Lab 10: OSD replacement~90 min · 6 objectives→
- Lab 11: MON recovery~90 min · 5 objectives→
- Lab 12: MDS recovery~60 min · 4 objectives→
- Lab 13: PG states and recovery~75 min · 5 objectives→
- Lab 14: Prometheus + Ceph MGR exporter~60 min · 5 objectives→
- Lab 15: Recovery tuning~60 min · 4 objectives→
- Lab 16: Capacity management~75 min · 4 objectives→
- Lab 17: RBD backup~60 min · 4 objectives→
- Lab 18: RBD with Proxmox~90 min · 4 objectives→
- Lab 19: RBD with Kubernetes~90 min · 5 objectives→
- Lab 20: CephFS with Kubernetes~75 min · 5 objectives→
- Lab 21: Network impairment~45 min · 5 objectives→
- Lab 22: CRUSH failure domain misconfiguration~60 min · 3 objectives→
- Lab 23: Capacity scenario~60 min · 4 objectives→
- Lab 24: CephFS with Proxmox CT~60 min · 3 objectives→
- Lab 25: Kubernetes PVC mount failure~45 min · 5 objectives→
- Lab 26: Proxmox VM storage latency diagnosis~60 min · 5 objectives→
- Lab 27: Upgrade in a staging cluster~180 min · 5 objectives→
- Lab 28: rados bench and rbd benchmark~60 min · 3 objectives→
- Lab 29: Deep scrub and repair~60 min · 4 objectives→
- Lab 30: CephFS snapshot and subvolume restore~60 min · 5 objectives→
Break/Fix scenarios
Troubleshooting drills. Each scenario gives symptoms and evidence, then hides the solution behind a reveal.
ceph-backfillfull
ceph-cephfs
ceph-crush-misplace
ceph-disk-health
ceph-mgr-failover
ceph-slow-ops
ceph-osd
- advancedMultiple OSD failure - exceeds expected tolerance~28 min · 3 symptoms→
- intermediateNVMe-oF target failure~30 min · 2 symptoms→
- intermediateOSD is down - degraded cluster, recovery in progress~20 min · 4 symptoms→
- advancedOSD fails repeatedly after restart - journal / DB device failure~25 min · 3 symptoms→
ceph-net-degraded
ceph-net-partition
ceph-pg-degraded
ceph-inconsistent-pg
ceph-rados
ceph-rbd
ceph-rbd-snap
ceph-upgrade-mixed
Assessments
Production-readiness examinations. Each combines auto-scored questions with scenario-based rubrics you can use to grade yourself.