Skip to main content
RunBook Academy

storage · security · cloud

Ceph & Distributed Storage for Production Sysadmins

A hands-on, fully visual course that takes a systems administrator from "I have heard of Ceph" to "I can take operational responsibility for a business-critical Ceph cluster on bare metal or in any VM platform." Covers the distributed-storage concepts underneath Ceph, the Ceph architecture (RADOS, MON, MGR, OSD, MDS, RGW), CRUSH placement, placement groups, replication and erasure coding, RBD, CephFS, RGW, cephadm deployment, capacity management, performance engineering, observability, security, backup and disaster recovery, Proxmox VE integration, Kubernetes CSI integration, and a production capstone.

Who this is for

  • Linux systems administrators taking operational responsibility for Ceph
  • Storage engineers running distributed storage at scale
  • Virtualisation engineers using Ceph with Proxmox VE
  • Platform engineers integrating Ceph with Kubernetes
  • SREs and DevOps engineers owning Ceph-backed services
  • Infrastructure engineers designing failure-domain-aware storage
  • Database engineers evaluating Ceph for database storage

Prerequisites

  • Comfortable on the Linux command line
  • Familiar with networking fundamentals (IP, VLAN, DNS, TLS)
  • Some exposure to storage devices, filesystems, and RAID/LVM
  • Comfort reading TypeScript, YAML, and shell output

Other RunBook Academy courses

  • Linux — recommended. Ceph runs on Linux and inherits every Linux primitive it depends on (systemd, networking, filesystems, performance tooling). The Linux course covers the depth this Ceph course then uses.
  • Proxmox VE — recommended. Proxmox VE is one of the most common consumers of Ceph as a hypervisor storage backend. The Proxmox course has a Part VIII Ceph deep dive; the two courses are complementary.
  • Kubernetes — recommended. Kubernetes consumes Ceph via CSI (csi-rbd and csi-cephfs). The Kubernetes course covers the workload, scheduler, and CSI mechanics this Ceph course then integrates with.
  • Observability — recommended. Prometheus, Grafana, Loki are the standard observability stack for production Ceph clusters. The Observability course covers the depth this Ceph course then integrates with.
  • Docker & Containers — recommended. cephadm deploys Ceph daemons as containers, and the Kubernetes-Ceph integration uses OCI artefacts. The Docker course covers the runtime primitives this Ceph course then relies on.
  • Ansible — recommended. Day-2 Ceph operations (configuration as code, fleet-wide state, recovery playbooks) benefit from automated configuration management.
  • Terraform — recommended. Where a Ceph cluster is provisioned by infrastructure-as-code rather than cephadm alone, Terraform (or OpenTofu) is a common approach.
  • VyOS — recommended. On-prem Ceph clusters depend on network design — bond/LACP, MTU, VLANs, BGP. The VyOS course covers the depth a Ceph designer assumes.

What you'll be able to do

After completing this course, you should be capable of independently:

  • Explain distributed storage concepts at the depth required to reason about Ceph under failure
  • Design a Ceph cluster failure-domain-aware: racks, hosts, power domains, network paths
  • Deploy a Ceph cluster with cephadm on bare metal or nested VMs
  • Configure and operate RADOS, MON, MGR, OSD, MDS, RGW with production discipline
  • Build and verify CRUSH rules that actually place replicas on separate failure domains
  • Operate replicated and erasure-coded pools, choosing between them per workload
  • Operate RBD, CephFS, and RGW with least-privilege cephx capabilities
  • Diagnose OSD failure, recovery, backfill, and the difference between them
  • Diagnose MON quorum loss, network partition, and CRUSH-misplaced replicas
  • Manage capacity, nearfull, backfillfull, full thresholds without losing the cluster
  • Tune recovery speed against client I/O using production values, not paper values
  • Investigate slow ops and blocked ops from OSD, client, and network evidence
  • Integrate Prometheus and Grafana for Ceph observability
  • Author alerts for quorum loss, OSD down, degraded PGs, nearfull, slow ops
  • Replace a failed OSD, a failed disk, and a failed host safely
  • Add and remove storage nodes while preserving durability and availability
  • Maintain CRUSH topology and reweight without losing placement guarantees
  • Upgrade a Ceph cluster following the documented cephadm path safely
  • Apply cluster maintenance flags correctly and remove them
  • Perform node maintenance and network maintenance without taking the cluster down
  • Harden cephx, network segmentation, dashboard access, and RGW user management
  • Rotate, back up, and recover cephx keys and dashboard credentials
  • Distinguish Ceph redundancy from backup, and design backup strategy accordingly
  • Back up RBD, CephFS, and RGW data with the appropriate RPO for each workload
  • Recover from multi-OSD, single-host, single-rack, and full-cluster failures
  • Integrate Ceph with Proxmox VE (hyper-converged and dedicated) and reason about trade-offs
  • Integrate Ceph with Kubernetes via csi-rbd and csi-cephfs
  • Make the architecture decision: when Ceph is the right answer and when it is not
  • Complete a capstone: a production 4-node Ceph cluster with RBD/CephFS/RGW, Proxmox and Kubernetes consumers, full observability, and 12 controlled-failure incidents

Curriculum overview

130 planned parts · 744 lessons currently published.

Part I

Storage Fundamentals

Block, file, object, DAS, SAN, NAS, distributed, persistent data, durability vs availability.

6 lessons

Part II

Storage Performance Fundamentals

IOPS, throughput, latency, queue depth, block size, random vs sequential I/O, read/write ratios, tail latency.

6 lessons

Part III

Storage Hardware

HDD, SATA SSD, enterprise SSD, NVMe, endurance, DWPD, latency, power-loss protection, SMART/NVMe health.

6 lessons

Part IV

Failure Domains

Disk, server, rack, power domain, datacentre, and why replica placement must align with actual physical failure domains.

6 lessons

Part V

Distributed Systems Foundations

Distributed state, node failure, network partition, consistency, availability, quorum, replication.

6 lessons

Part VI

Ceph Architecture

Architecture overview: clients, RADOS, MON/MGR/MDS, RADOS cluster, OSDs; the system underneath the CLI.

6 lessons

Part VII

RADOS

The reliable autonomic distributed object store: objects, pools, placement groups, CRUSH, replication, OSDs.

6 lessons

Part VIII

Monitors

MON responsibilities, cluster maps, quorum, Paxos/consensus concepts, monitor stores, election.

6 lessons

Part IX

Monitor Quorum

3 MONs, 2 required for quorum. More complex scenarios and quorum failure exercises.

6 lessons

Part X

Manager Daemons

MGR role, modules, dashboards, telemetry, orchestration integration, MGR vs MON.

6 lessons

Part XI

OSD Architecture

OSD daemon, device, BlueStore, memory, networking, heartbeat, PG ownership.

6 lessons

Part XII

BlueStore

Application write to RADOS to OSD to BlueStore to block device. DB, WAL, RocksDB, allocation, checksums.

6 lessons

Part XIII

CRUSH Fundamentals

Deterministic placement, buckets, roots, hosts, racks, device classes, rules.

6 lessons

Part XIV

CRUSH Failure Domains

Bucket hierarchy examples, replica placement mechanics, non-colocation enforcement.

6 lessons

Part XV

CRUSH Maps and Rules

How placement rules affect media class, replication, resilience, performance.

6 lessons

Part XVI

Device Classes

HDD, SSD, NVMe classes, custom placement strategies for mixed-storage clusters.

6 lessons

Part XVII

Pools

Pool purpose, replicated pools, erasure-coded pools, application tagging, quotas, autoscaling.

6 lessons

Part XVIII

Placement Groups

Object to hash to placement group to CRUSH to OSDs. Why PGs exist.

6 lessons

Part XIX

PG States

active, clean, degraded, undersized, peering, inactive, stale, inconsistent. Real diagnostic examples.

6 lessons

Part XX

PG Peering

What happens after OSD restart, OSD failure, cluster topology change. Acting set, up set, primary OSD.

6 lessons

Part XXI

PG Autoscale

PG autoscaler behaviour, target size, workload changes, recommendations.

6 lessons

Part XXII

PG Investigation

Reading PG state, transitioning through states, deciding intervention.

6 lessons

Part XXIII

Replication

size and min_size semantics. Normal operation, degraded operation, write availability, durability.

6 lessons

Part XXIV

Replica Failure Scenarios

Two replicas available, one failed. Three replicas, two failed. Reasoning about safety.

6 lessons

Part XXV

Erasure Coding Fundamentals

k data chunks plus m parity chunks. Capacity efficiency and failure tolerance.

6 lessons

Part XXVI

Erasure Coding Trade-offs

CPU overhead, write amplification, small writes, recovery, durability, storage efficiency.

6 lessons

Part XXVII

Replication vs Erasure Coding

Architecture decision exercises for VM disks, object archives, databases.

6 lessons

Part XXVIII

Ceph Networking

Public network, cluster network where relevant, client, replication, recovery, heartbeat, MTU, bonding, VLANs.

6 lessons

Part XXIX

Network Design

10/25/40/100 GbE, redundancy, switch failure domains, LACP, routed designs.

6 lessons

Part XXX

Network Failure Behaviour

Packet loss, latency, asymmetric paths, switch failure, bond member failure. Storage symptoms.

6 lessons

Part XXXI

Ceph Authentication (cephx)

cephx, entities, keyrings, capabilities.

6 lessons

Part XXXII

Least Privilege Capabilities

Capability design for RBD clients, CephFS, RGW, monitoring, administration.

6 lessons

Part XXXIII

Encryption

Network encryption where supported, disk and device encryption, application-level encryption, RGW SSE.

6 lessons

Part XXXIV

Multi-Tenancy Concepts

Isolation across RBD pools, CephFS subvolumes, RGW users and tenants.

6 lessons

Part XXXV

RBD Architecture

VM / host to RBD client to RADOS to pool to PG to OSDs.

6 lessons

Part XXXVI

RBD Images

Images, features, snapshots, clones, flattening, resizing.

6 lessons

Part XXXVII

RBD Snapshots

Snapshots are not backup. Same-cluster dependency, failure domains, rollback implications.

6 lessons

Part XXXVIII

RBD Performance

Queue depth, caching, workload block size, replication overhead, network, OSD latency.

6 lessons

Part XXXIX

RBD Troubleshooting

Image unavailable, client auth, mapping, latency, blocked requests, client and kernel version compatibility.

6 lessons

Part XL

CephFS Architecture

Client to MDS to metadata to RADOS data pools. Metadata vs file contents.

6 lessons

Part XLI

Metadata Servers (MDS)

Active, standby, metadata cache, failover. MDS rank 0, recovery, session behaviour.

6 lessons

Part XLII

CephFS Operations

Filesystem creation, clients, quotas, snapshots, subvolumes.

6 lessons

Part XLIII

CephFS Failure Scenarios

MDS failure, slow metadata operations, full pools, client issues, session timeouts.

6 lessons

Part XLIV

Object Storage Foundations

Objects, buckets, S3 concepts, object namespaces, versioning, lifecycle.

6 lessons

Part XLV

RADOS Gateway (RGW)

S3 client to RGW to RADOS architecture, daemon placement, multisite concepts.

6 lessons

Part XLVI

RGW Users and Credentials

Access keys, secret keys, users, subusers, quotas, bucket policies.

6 lessons

Part XLVII

RGW High Availability

Multiple gateways, load balancing, zonegroup / zone placement, eventual consistency in multisite.

6 lessons

Part XLVIII

RGW Troubleshooting

Authentication, bucket access, backend pool, latency, load balancer.

6 lessons

Part XLIX

cephadm

Bootstrap, hosts, services, daemons, inventory, labels, the cephadm orchestrator.

6 lessons

Part L

Cluster Deployment

Production deployment guidance. Node count, MON placement, MGR, OSD layout, networks, time, DNS, storage devices.

6 lessons

Part LI

Host Preparation

Linux, networking, time sync, hostname / DNS, disks, container runtime, partitions, device-by-id.

6 lessons

Part LII

Time Synchronisation

Time as a production dependency. Consequences for distributed systems, MON election, recovery, log correlation.

6 lessons

Part LIII

Cluster Health

HEALTH_OK, HEALTH_WARN, HEALTH_ERR. Reading health messages, not just health states.

6 lessons

Part LIV

ceph status and health detail

ceph -s, ceph health detail, ceph health with output interpretation.

6 lessons

Part LV

OSD States

up / down and in / out. Why these are separate dimensions. Critical lesson.

6 lessons

Part LVI

OSD Failure

OSD down to PG degraded to recovery decisions to OSD out to backfill. The full failure cascade.

6 lessons

Part LVII

Replacing Failed OSDs

Identify disk, verify failure, safely remove, replace, recreate, monitor recovery.

6 lessons

Part LVIII

Recovery

Recovery semantics, missing replicas, object reconstruction, recovery priorities, impact on production I/O.

6 lessons

Part LIX

Backfill

Why topology and capacity changes trigger data movement. Difference between recovery and backfill.

6 lessons

Part LX

Recovery Tuning

Balancing client I/O vs recovery speed. osd_recovery_max_active, osd_recovery_sleep, osd_max_backfills.

6 lessons

Part LXI

Scrubbing

Scrub vs deep scrub. Checksums, consistency, daily and weekly schedules, scrubbing performance impact.

6 lessons

Part LXII

Inconsistent PGs

How inconsistent objects are detected and repaired. Strong warnings around pg repair.

6 lessons

Part LXIII

Capacity Management

Raw, usable, replication overhead, EC overhead, reserved headroom, recovery capacity, backfill capacity.

6 lessons

Part LXIV

Nearfull, Backfillfull and Full

Current threshold behaviour. Why clusters need free capacity to recover.

6 lessons

Part LXV

Why Full Clusters Are Dangerous

Cluster becomes full, writes blocked, recovery constrained, operational options shrink.

6 lessons

Part LXVI

Capacity Forecasting

Growth rate, retention, workload expansion, recovery headroom, node-addition lead time.

6 lessons

Part LXVII

Performance Methodology

Application to client to network to OSD to device. Systematic isolation.

6 lessons

Part LXVIII

OSD Latency

Identifying slow OSDs, correlating ceph daemonperf, perf dump, OSD latency histograms.

6 lessons

Part LXIX

Disk Performance

Disk-side latency, queueing, device errors, SMART / NVMe health integration with Ceph.

6 lessons

Part LXX

Network Performance

Storage latency can actually be network latency. Bonding, MTU, asymmetric routes, switch load.

6 lessons

Part LXXI

Client Performance

RBD client behaviour, queue depth, concurrency, caching, krbd vs librbd vs rbd-nbd.

6 lessons

Part LXXII

Benchmarking

rados bench, rbd bench, fio against RBD, ceph bench where appropriate. Never against production.

6 lessons

Part LXXIII

Benchmark Interpretation

Latency distribution, IOPS, throughput, queue depth, unrealistic benchmarks, workload mismatch.

6 lessons

Part LXXIV

Observability

Monitoring cluster health, MON quorum, OSD states, PG states, capacity, latency, recovery, network, device health.

6 lessons

Part LXXV

Prometheus Metrics

MGR Prometheus exporter, high-value metrics, histograms vs counters, latency observation.

6 lessons

Part LXXVI

Grafana Dashboards

Per-cluster health, per-pool IOPS, per-OSD latency, recovery, capacity, network.

6 lessons

Part LXXVII

Alerting

Actionable alerts: MON quorum loss, OSD down, degraded PGs, inconsistent PG, nearfull, slow ops.

6 lessons

Part LXXVIII

Monitoring the Monitoring Path

What happens if Ceph is failing and the telemetry pipeline is also unavailable.

6 lessons

Part LXXIX

Slow Ops

What slow ops mean. OSD, client, network causes. Correlating evidence across slow ops and daemonperf.

6 lessons

Part LXXX

Blocked Operations

blocked vs slow ops. Diagnostic interpretation.

6 lessons

Part LXXXI

Proxmox Integration

RBD storage, CephFS where applicable, hyper-converged design, dedicated storage clusters, networking.

6 lessons

Part LXXXII

Hyper-Converged Ceph

Compute plus Ceph OSD on same nodes. Cost, performance contention, failure domains, maintenance.

6 lessons

Part LXXXIII

Dedicated Ceph Cluster

Dedicated storage nodes vs hyper-converged. Trade-offs.

6 lessons

Part LXXXIV

Proxmox Failure Scenarios

VM I/O latency, failed OSD, node maintenance, quorum dependencies.

6 lessons

Part LXXXV

Kubernetes Integration

CSI (csi-rbd, csi-cephfs). StorageClass, PVC, secrets, provisioner, snapshotter, topology awareness.

6 lessons

Part LXXXVI

Kubernetes RBD

Block volumes via csi-rbd. StorageClass parameters, RBD image format, topology, dynamic provisioning.

6 lessons

Part LXXXVII

Kubernetes CephFS

Shared filesystem via csi-cephfs. Subvolume provisioning, subvolume snapshots, dynamic provisioning.

6 lessons

Part LXXXVIII

Kubernetes Storage Failure Scenarios

PVC mount failure, CSI problems, Ceph auth, pool failure, backend capacity, stale snapshot.

6 lessons

Part LXXXIX

Rook Concepts

Rook as a Kubernetes operator for Ceph. Forward-pointer, awareness lesson, not a deep dive.

6 lessons

Part XC

Scaling Out

Adding OSDs, hosts, capacity, MONs only when justified. Resulting rebalancing.

6 lessons

Part XCI

Adding Storage Nodes

Pre-checks: failure domain, network, capacity, CRUSH placement, recovery impact.

6 lessons

Part XCII

Removing Storage Nodes

Safe evacuation and removal. Strong warnings against simply deleting OSDs.

6 lessons

Part XCIII

Changing CRUSH Topology

Blast radius, data movement impact, crush reweight, crush reweight-all, crush move.

6 lessons

Part XCIV

Hardware Replacement

Failed disk, failing disk, failed server, replacement node.

6 lessons

Part XCV

Maintenance Flags

noout, norebalance, norecover, nobackfill, nodeep-scrub. Strong warning: never leave enabled indefinitely.

6 lessons

Part XCVI

Node Maintenance

Planned reboot, kernel update, OSD host maintenance, daemon restarts.

6 lessons

Part XCVII

Network Maintenance

Switch / link maintenance with storage redundancy, bond member replacement, MTU changes.

6 lessons

Part XCVIII

Software Upgrades

Safe Ceph upgrade workflows. Containerised cephadm upgrade path.

6 lessons

Part XCIX

Upgrade Planning

Read release notes, check health, backup config/state, validate version path, upgrade staging, upgrade cluster, monitor.

6 lessons

Part C

Upgrade Failure Recovery

Scenarios: daemon fails after upgrade, mixed-version state, module/plugin incompatibility, MGR failover.

6 lessons

Part CI

Security Hardening

cephx, admin access, network segmentation, keyrings, least privilege, encrypted networks, host hardening.

6 lessons

Part CII

Management Security

cephadm SSH, admin keys, dashboard access, monitoring access, REST API authentication, mgr / balancer modules.

6 lessons

Part CIII

Secrets and Key Management

Keyring protection, client keys, rotation, backup of authentication state, mvcc and metadata db.

6 lessons

Part CIV

Multi-Tenancy in Practice

RBD pool-isolation, CephFS subvolume-isolation, RGW users and tenants, capability scoping.

6 lessons

Part CV

Backup Strategy

Ceph redundancy is NOT backup. Replication protects against device and node failure. Backup protects against deletion, corruption, ransomware, operator error.

6 lessons

Part CVI

RBD Backup

RBD snapshots, rbd export, rbd export-diff, incremental strategies, application consistency.

6 lessons

Part CVII

CephFS Backup

CephFS snapshots, file-level backups, subvolume snapshots, recovery validation.

6 lessons

Part CVIII

RGW Backup and Replication

Object storage backup, RGW multisite replication, RPO and RTO modelling.

6 lessons

Part CIX

Disaster Recovery

Multiple OSD loss, MON loss, site loss, complete cluster loss. The recovery scenario set.

6 lessons

Part CX

Monitor Recovery

Rebuilding and recovering MONs safely. monmap, monitor store recovery, hostname changes.

6 lessons

Part CXI

Manager Recovery

Active / standby MGR, recover MGR services.

6 lessons

Part CXII

OSD Host Loss

Entire OSD server disappears. Reasoning about CRUSH failure domain and data safety.

6 lessons

Part CXIII

Multiple OSD Failure

Replication size 3 plus multiple OSD losses. Reasoning about remaining replicas.

6 lessons

Part CXIV

Complete Storage Node Loss

Recovering safely after a host is gone.

6 lessons

Part CXV

Cluster-Wide Capacity Incident

nearfull, backfillfull, full. Immediate priorities.

6 lessons

Part CXVI

Network Partition

Ceph behaviour during partial network connectivity. MON quorum and OSD communication implications.

6 lessons

Part CXVII

Lost Monitor Quorum

Deep troubleshooting for MON quorum loss, including split-brain recovery.

6 lessons

Part CXVIII

Data Integrity Incident

Inconsistent PGs, checksum errors, disk errors, repair decisions.

6 lessons

Part CXIX

Disaster Recovery Architecture

Primary Ceph cluster to backups and replication to independent recovery domain.

6 lessons

Part CXX

Multi-Site Concepts

RGW multisite, asynchronous replication concepts, RPO, RTO.

6 lessons

Part CXXI

Storage Architecture Decision-Making

When Ceph is appropriate and when it is not. Workload fit, operational capacity, scale, latency.

6 lessons

Part CXXII

Small Cluster Risks

3-node Ceph cluster trade-offs. Recovery bandwidth, maintenance headroom, capacity, fault tolerance.

6 lessons

Part CXXIII

Capacity and Failure Planning

Normal plus growth plus node failure plus recovery headroom plus maintenance headroom.

6 lessons

Part CXXIV

Production Reference Architecture

The complete reference design: 6+ nodes, MON/MGR placement, OSD layout, network architecture, services, monitoring, automation, backups.

6 lessons

Part Labs

Labs

Hands-on disposable-cluster labs (cephadm build, OSD lifecycle, RBD, CephFS, RGW, recovery, capacity, Proxmox, Kubernetes CSI, network impairment).

0 lessons

Part Runbooks

Runbooks

Operational procedures for build, OSD replacement, MON recovery, recovery tuning, capacity incidents, DR, network maintenance.

0 lessons

Part Checklists

Checklists

Production readiness, hardware readiness, network readiness, CRUSH review, pool design, capacity planning, security, RBD/CephFS/RGW readiness, observability, OSD replacement, node maintenance, pre-upgrade, post-upgrade, backup DR, Proxmox + Ceph readiness, Kubernetes + Ceph readiness.

0 lessons

Part Breakfix

Break/Fix Scenarios

Evidence-first diagnosis of Ceph failure modes (OSD, MON, MDS, RBD, CephFS, RGW, network, capacity, recovery, upgrades).

0 lessons

Part Capstone

Capstone: Mission-Critical Ceph

A complete production Ceph estate end to end: 4-node cluster with RBD/CephFS/RGW, Proxmox and Kubernetes consumers, full observability, backup, and 12 controlled-failure incidents.

0 lessons

Part Final

Final Assessment

Final theory assessment plus final practical assessment of an inherited Ceph estate.

0 lessons

Verified against

  • CephvTentacle 20.2.x· released 2025-11-18· verified 2026-08-17
  • CephvSquid 19.2.x (supported previous)· released 2024-09-26· verified 2026-08-17
  • cephadmvmatches the verified Ceph release· verified 2026-08-17
  • podmanv4.x· verified 2026-08-17
  • csi-rbd and csi-cephfsvcurrent· verified 2026-08-17
  • RBD / CephFS / RGWvcurrent (matches Ceph release)· verified 2026-08-17
  • Linux kernelv5.15+ (5.10 minimum)· verified 2026-08-17
  • Ubuntuv24.04 LTS (Ceph host baseline)· verified 2026-08-17
  • Debianv12 (Bookworm) (Ceph host baseline)· verified 2026-08-17
  • Rocky Linux / RHEL / AlmaLinuxv9.x (Ceph host baseline)· verified 2026-08-17
  • Proxmox VEv9.x (cross-course integration)· verified 2026-08-17
  • Kubernetesv1.31+ (cross-course integration)· verified 2026-08-17