Skip to main content
RunBook Academy

Ceph · Curriculum

Curriculum

744 lessons across 130 parts. Lessons build on each other; later parts assume familiarity with earlier material.

Part I

Storage Fundamentals

Block, file, object, DAS, SAN, NAS, distributed, persistent data, durability vs availability.

6 lessons
  1. 01Storage primitives — why the primitive is the contractStorage Fundamentals · foundation · ~14 min
  2. 01Block storage - the primitive every database depends onStorage primitives · foundation · ~18 min
  3. 02File storage - the shared namespace primitiveStorage primitives · foundation · ~16 min
  4. 04Object storage — the flat namespace primitiveStorage Fundamentals · foundation · ~15 min
  5. 05DAS, SAN, NAS — topology is not the same as primitiveStorage Fundamentals · foundation · ~14 min
  6. 06Durability and availability — two different numbersStorage Fundamentals · intermediate · ~16 min

Part II

Storage Performance Fundamentals

IOPS, throughput, latency, queue depth, block size, random vs sequential I/O, read/write ratios, tail latency.

6 lessons
  1. 01IOPS, throughput, latency — and why two of them fightStorage Performance Fundamentals · foundation · ~15 min
  2. 02Queue depth and block size — the two knobs that set the costStorage Performance Fundamentals · intermediate · ~15 min
  3. 03Random and sequential, read and write — the four profilesStorage Performance Fundamentals · foundation · ~14 min
  4. 04Tail latency — why p99 is the number that mattersStorage Performance Fundamentals · intermediate · ~15 min
  5. 05High throughput is not low latencyStorage Performance Fundamentals · intermediate · ~16 min
  6. 06Fast disks are not fast distributed storageStorage Performance Fundamentals · intermediate · ~15 min

Part III

Storage Hardware

HDD, SATA SSD, enterprise SSD, NVMe, endurance, DWPD, latency, power-loss protection, SMART/NVMe health.

6 lessons
  1. 01HDD — where spinning disks still belong in a Ceph clusterStorage Hardware · foundation · ~14 min
  2. 02SATA SSD — the general-purpose Ceph OSDStorage Hardware · foundation · ~13 min
  3. 03Enterprise SSD — power-loss protection and why it is not optionalStorage Hardware · intermediate · ~15 min
  4. 04NVMe — parallelism, not just speedStorage Hardware · intermediate · ~15 min
  5. 05Endurance and DWPD — sizing flash for years of writesStorage Hardware · intermediate · ~15 min
  6. 06Device health — SMART, NVMe logs, and what actually predicts failureStorage Hardware · intermediate · ~14 min

Part IV

Failure Domains

Disk, server, rack, power domain, datacentre, and why replica placement must align with actual physical failure domains.

6 lessons
  1. 01What a failure domain is, and why the answer is never "the disk"Failure Domains · foundation · ~14 min
  2. 02Physical topology — mapping racks, rows, and chassis into CRUSHFailure Domains · intermediate · ~15 min
  3. 03Power domains — the failure boundary CRUSH cannot seeFailure Domains · intermediate · ~14 min
  4. 04Network domains — switches, uplinks, and partition behaviourFailure Domains · intermediate · ~15 min
  5. 05Correlated failure — why independent probability estimates misleadFailure Domains · advanced · ~15 min
  6. 06Placement decisions — turning failure domains into pool configurationFailure Domains · intermediate · ~16 min

Part V

Distributed Systems Foundations

Distributed state, node failure, network partition, consistency, availability, quorum, replication.

6 lessons
  1. 01Distributed state — where the truth livesDistributed Systems Foundations · foundation · ~14 min
  2. 02Node failure — "did not answer in time" is the only signal you getDistributed Systems Foundations · foundation · ~15 min
  3. 03Network partition — when both halves are alive and neither is sureDistributed Systems Foundations · intermediate · ~15 min
  4. 04CAP in operational terms — what Ceph gives up and whenDistributed Systems Foundations · intermediate · ~14 min
  5. 05Quorum — why majority is the only safe ruleDistributed Systems Foundations · foundation · ~14 min
  6. 06Replication — the cost of surviving a node failureDistributed Systems Foundations · foundation · ~15 min

Part VI

Ceph Architecture

Architecture overview: clients, RADOS, MON/MGR/MDS, RADOS cluster, OSDs; the system underneath the CLI.

6 lessons
  1. 01Ceph architecture overview — the path a write takesCeph Architecture · foundation · ~16 min
  2. 02The storage layer cake — RBD, CephFS, and RGW on one RADOSCeph Architecture · foundation · ~15 min
  3. 03The daemons — mon, mgr, osd, mds, rgw and what each one ownsCeph Architecture · foundation · ~15 min
  4. 04Cluster map and cluster identity — FSID, monmap, osdmapCeph Architecture · intermediate · ~15 min
  5. 05The client protocol and msgr2 — how clients talk to the clusterCeph Architecture · intermediate · ~15 min
  6. 06Deployment models — cephadm and what it replacedCeph Architecture · intermediate · ~15 min

Part VII

RADOS

The reliable autonomic distributed object store: objects, pools, placement groups, CRUSH, replication, OSDs.

6 lessons
  1. 01Objects — the unit of storage everything else is built fromRADOS · foundation · ~15 min
  2. 02Pools — the unit of policyRADOS · foundation · ~15 min
  3. 03CRUSH — placement without a lookup tableRADOS · intermediate · ~16 min
  4. 04Replication in RADOS — primary, acting set, and the commitRADOS · intermediate · ~15 min
  5. 05Failure recovery in RADOS — what happens after an OSD goes awayRADOS · intermediate · ~16 min
  6. 06librados — the API every Ceph service is built onRADOS · advanced · ~15 min

Part VIII

Monitors

MON responsibilities, cluster maps, quorum, Paxos/consensus concepts, monitor stores, election.

6 lessons
  1. 01What the monitors actually doMonitors · foundation · ~15 min
  2. 02Cluster maps in detail — what each one holds and who reads itMonitors · intermediate · ~15 min
  3. 03Paxos at the depth a Ceph admin needsMonitors · advanced · ~15 min
  4. 04The monitor store — RocksDB, growth, and backupsMonitors · advanced · ~16 min
  5. 05Monitor election and the leader leaseMonitors · advanced · ~15 min
  6. 06Why three monitors — and why not two, four, or sevenMonitors · foundation · ~14 min

Part IX

Monitor Quorum

3 MONs, 2 required for quorum. More complex scenarios and quorum failure exercises.

6 lessons
  1. 01The quorum model in operationMonitor Quorum · intermediate · ~15 min
  2. 02Losing one monitor — a stable state, not an emergencyMonitor Quorum · intermediate · ~15 min
  3. 03Losing two monitors — the cluster stopsMonitor Quorum · advanced · ~16 min
  4. 04Asymmetric failures — when reachability is not mutualMonitor Quorum · advanced · ~16 min
  5. 05Monitor recovery order — what to try, in what sequenceMonitor Quorum · advanced · ~16 min
  6. 06Quorum failure exercises — two scenarios worth rehearsingMonitor Quorum · advanced · ~17 min

Part X

Manager Daemons

MGR role, modules, dashboards, telemetry, orchestration integration, MGR vs MON.

6 lessons
  1. 01The manager daemon — what it does and what it does notManager Daemons · foundation · ~14 min
  2. 02Manager modules — which to enable and which to leave offManager Daemons · intermediate · ~15 min
  3. 03The dashboard — useful, and an exposed web serviceManager Daemons · intermediate · ~15 min
  4. 04The orchestrator — how ceph orch drives cephadmManager Daemons · intermediate · ~16 min
  5. 05Manager failover — what happens and what to verifyManager Daemons · intermediate · ~14 min
  6. 06The balancer — evening out PG distribution with upmapManager Daemons · intermediate · ~16 min

Part XI

OSD Architecture

OSD daemon, device, BlueStore, memory, networking, heartbeat, PG ownership.

6 lessons
  1. 01The OSD process — one daemon, many responsibilitiesOSD Architecture · foundation · ~15 min
  2. 02BlueStore is the only backend — what that meansOSD Architecture · foundation · ~15 min
  3. 03OSDs and devices — one, or several, per driveOSD Architecture · intermediate · ~15 min
  4. 04OSD memory and threads — where the resources goOSD Architecture · advanced · ~16 min
  5. 05Heartbeats — how OSDs detect each other and report failureOSD Architecture · intermediate · ~15 min
  6. 06PG ownership — primaries, replicas, and why it mattersOSD Architecture · advanced · ~16 min

Part XII

BlueStore

Application write to RADOS to OSD to BlueStore to block device. DB, WAL, RocksDB, allocation, checksums.

6 lessons
  1. 01BlueFS, DB, and WAL — the three pieces inside BlueStoreBlueStore · intermediate · ~16 min
  2. 02The write-ahead log — durability ordering for metadataBlueStore · advanced · ~15 min
  3. 03The DB device — RocksDB, omap, and compactionBlueStore · advanced · ~16 min
  4. 04Allocation — how BlueStore lays data on the deviceBlueStore · advanced · ~16 min
  5. 05Checksums — the promise that data comes back as it went inBlueStore · intermediate · ~15 min
  6. 06Collections — how BlueStore organises objects by placement groupBlueStore · advanced · ~15 min

Part XIII

CRUSH Fundamentals

Deterministic placement, buckets, roots, hosts, racks, device classes, rules.

6 lessons
  1. 01The CRUSH algorithm — computing placement from the mapCRUSH Fundamentals · intermediate · ~16 min
  2. 02The CRUSH map — buckets, items, and weightsCRUSH Fundamentals · intermediate · ~16 min
  3. 03The select operation — choosing items at each levelCRUSH Fundamentals · advanced · ~16 min
  4. 04The take operation — where the descent beginsCRUSH Fundamentals · intermediate · ~15 min
  5. 05The emit operation and reading a complete ruleCRUSH Fundamentals · intermediate · ~15 min
  6. 06CRUSH versus centralised lookup — what the trade actually isCRUSH Fundamentals · advanced · ~15 min

Part XIV

CRUSH Failure Domains

Bucket hierarchy examples, replica placement mechanics, non-colocation enforcement.

6 lessons
  1. 01Host as the default failure domainCRUSH Failure Domains · foundation · ~15 min
  2. 02Rack failure domain — aligning replicas with the buildingCRUSH Failure Domains · intermediate · ~16 min
  3. 03Chassis, room, and row — extending the hierarchyCRUSH Failure Domains · intermediate · ~15 min
  4. 04Weights across failure domainsCRUSH Failure Domains · intermediate · ~16 min
  5. 05What changes when weights changeCRUSH Failure Domains · intermediate · ~16 min
  6. 06When CRUSH cannot satisfy the ruleCRUSH Failure Domains · intermediate · ~16 min

Part XV

CRUSH Maps and Rules

How placement rules affect media class, replication, resilience, performance.

6 lessons
  1. 01The replicated rule — the form you will write most oftenCRUSH Maps and Rules · foundation · ~15 min
  2. 02The erasure-coded rule — indep, chunk positions, and profilesCRUSH Maps and Rules · advanced · ~16 min
  3. 03Multiple rules on one clusterCRUSH Maps and Rules · intermediate · ~15 min
  4. 04Device-class rules — one cluster, several performance tiersCRUSH Maps and Rules · intermediate · ~15 min
  5. 05Editing rules — which changes are safe and which are notCRUSH Maps and Rules · advanced · ~16 min
  6. 06Rule and weight changes that move data — planning the whole eventCRUSH Maps and Rules · advanced · ~17 min

Part XVI

Device Classes

HDD, SSD, NVMe classes, custom placement strategies for mixed-storage clusters.

6 lessons
  1. 01How an OSD gets its device classDevice Classes · foundation · ~15 min
  2. 02CRUSH rules that select by classDevice Classes · intermediate · ~15 min
  3. 03Mixed-storage clusters — safe when separated, dangerous when notDevice Classes · intermediate · ~16 min
  4. 04The lifecycle of a device class assignmentDevice Classes · intermediate · ~15 min
  5. 05Wrong class placement — the most common pool mistakeDevice Classes · intermediate · ~15 min
  6. 06The performance impact of device class — the biggest lever you haveDevice Classes · intermediate · ~16 min

Part XVII

Pools

Pool purpose, replicated pools, erasure-coded pools, application tagging, quotas, autoscaling.

6 lessons
  1. 01Pool purpose — the unit where policy is decidedPools · foundation · ~15 min
  2. 02Replicated pools in productionPools · foundation · ~15 min
  3. 03Erasure-coded pools in productionPools · advanced · ~17 min
  4. 04Application tagging — telling Ceph who owns a poolPools · foundation · ~14 min
  5. 05Pool quotas — bounding capacity per poolPools · intermediate · ~15 min
  6. 06Pool autoscaling — letting the manager size pg_numPools · intermediate · ~16 min

Part XVIII

Placement Groups

Object to hash to placement group to CRUSH to OSDs. Why PGs exist.

6 lessons
  1. 01Why placement groups existPlacement Groups · foundation · ~15 min
  2. 02Object to PG — the hashing stepPlacement Groups · intermediate · ~15 min
  3. 03PG to OSD — the CRUSH stepPlacement Groups · intermediate · ~15 min
  4. 04PG count tuning — the numbers and what they costPlacement Groups · intermediate · ~16 min
  5. 05PG states — the vocabulary you read during every incidentPlacement Groups · intermediate · ~16 min
  6. 06PGs per pool — managing the cluster-wide totalPlacement Groups · intermediate · ~15 min

Part XIX

PG States

active, clean, degraded, undersized, peering, inactive, stale, inconsistent. Real diagnostic examples.

6 lessons
  1. 01active+clean — what the steady state actually guaranteesPG States · foundation · ~14 min
  2. 02active+degraded — serving with fewer copies than promisedPG States · intermediate · ~15 min
  3. 03active+undersized — the acting set is shortPG States · intermediate · ~15 min
  4. 04Peering — agreeing on what the PG containsPG States · advanced · ~16 min
  5. 05Inactive PGs — the most urgent state in CephPG States · expert · ~17 min
  6. 06Inconsistent PGs — scrub found replicas that disagreePG States · advanced · ~17 min

Part XX

PG Peering

What happens after OSD restart, OSD failure, cluster topology change. Acting set, up set, primary OSD.

6 lessons
  1. 01Up set, acting set, and primaryPG Peering · intermediate · ~15 min
  2. 02Peering history — finding the authoritative copyPG Peering · advanced · ~16 min
  3. 03Peering after an OSD restartPG Peering · intermediate · ~15 min
  4. 04Peering after an OSD failurePG Peering · intermediate · ~15 min
  5. 05Peering after a topology changePG Peering · advanced · ~16 min
  6. 06Stuck peering — diagnosis and resolutionPG Peering · expert · ~17 min

Part XXI

PG Autoscale

PG autoscaler behaviour, target size, workload changes, recommendations.

6 lessons
  1. 01pg_autoscale_mode: on, warn, and offPG Autoscale · intermediate · ~14 min
  2. 02Reading autoscaler suggestionsPG Autoscale · intermediate · ~16 min
  3. 03target_size_bytes and target_size_ratioPG Autoscale · intermediate · ~15 min
  4. 04How the autoscaler responds to workload changePG Autoscale · intermediate · ~16 min
  5. 05Applying autoscaler recommendations safelyPG Autoscale · advanced · ~18 min
  6. 06Bounding what the autoscaler may doPG Autoscale · advanced · ~17 min

Part XXII

PG Investigation

Reading PG state, transitioning through states, deciding intervention.

6 lessons
  1. 01Reading a single PG with pg queryPG Investigation · advanced · ~18 min
  2. 02Cluster-wide PG state: pg stat, pg dump, health detailPG Investigation · intermediate · ~16 min
  3. 03When to intervene, and when to waitPG Investigation · advanced · ~17 min
  4. 04PGs stuck scrubbing or deep-scrubbingPG Investigation · advanced · ~16 min
  5. 05PGs stuck in recoveryPG Investigation · advanced · ~18 min
  6. 06Deep investigation: per-OSD latency and logsPG Investigation · expert · ~20 min

Part XXIII

Replication

size and min_size semantics. Normal operation, degraded operation, write availability, durability.

6 lessons
  1. 01size and min_size: what each one controlsReplication · foundation · ~15 min
  2. 02How a replicated write actually completesReplication · intermediate · ~16 min
  3. 03Writes while a replica is missingReplication · intermediate · ~17 min
  4. 04Reads on a degraded poolReplication · intermediate · ~15 min
  5. 05Choosing size: what each value actually buysReplication · advanced · ~18 min
  6. 06Why min_size 1 loses dataReplication · advanced · ~17 min

Part XXIV

Replica Failure Scenarios

Two replicas available, one failed. Three replicas, two failed. Reasoning about safety.

6 lessons
  1. 01One OSD fails on a size-3 poolReplica Failure Scenarios · intermediate · ~16 min
  2. 02Two OSDs fail on a size-3 poolReplica Failure Scenarios · advanced · ~17 min
  3. 03Cascading and correlated failureReplica Failure Scenarios · advanced · ~18 min
  4. 04Reasoning about data safety from the cluster mapReplica Failure Scenarios · advanced · ~17 min
  5. 05The acting set is the safety boundaryReplica Failure Scenarios · intermediate · ~16 min
  6. 06Recovery restores redundancy; it does not change the boundaryReplica Failure Scenarios · advanced · ~17 min

Part XXV

Erasure Coding Fundamentals

k data chunks plus m parity chunks. Capacity efficiency and failure tolerance.

6 lessons
  1. 01k and m: what an erasure code profile meansErasure Coding Fundamentals · intermediate · ~17 min
  2. 02Capacity efficiency and the real cost of a profileErasure Coding Fundamentals · intermediate · ~17 min
  3. 03How EC chunks are placedErasure Coding Fundamentals · advanced · ~18 min
  4. 04EC recovery and what it costsErasure Coding Fundamentals · advanced · ~18 min
  5. 05Choosing among the common profilesErasure Coding Fundamentals · advanced · ~18 min
  6. 06EC for RBD and CephFS: what it takes and when it fitsErasure Coding Fundamentals · advanced · ~18 min

Part XXVI

Erasure Coding Trade-offs

CPU overhead, write amplification, small writes, recovery, durability, storage efficiency.

6 lessons
  1. 01CPU cost of erasure codingErasure Coding Trade-offs · advanced · ~16 min
  2. 02Write amplification on EC poolsErasure Coding Trade-offs · advanced · ~17 min
  3. 03Why EC handles small writes badlyErasure Coding Trade-offs · advanced · ~16 min
  4. 04The recovery cost of erasure codingErasure Coding Trade-offs · advanced · ~17 min
  5. 05EC durability compared with replicationErasure Coding Trade-offs · expert · ~18 min
  6. 06Capacity planning for erasure-coded poolsErasure Coding Trade-offs · advanced · ~18 min

Part XXVII

Replication vs Erasure Coding

Architecture decision exercises for VM disks, object archives, databases.

6 lessons
  1. 01Why RBD generally wants replicationReplication vs Erasure Coding · intermediate · ~16 min
  2. 02Object archives are where EC paysReplication vs Erasure Coding · intermediate · ~17 min
  3. 03Large immutable data: the clearest EC caseReplication vs Erasure Coding · intermediate · ~16 min
  4. 04Databases: primary data replicated, everything around it ECReplication vs Erasure Coding · advanced · ~17 min
  5. 05A decision framework for pool typeReplication vs Erasure Coding · advanced · ~17 min
  6. 06Running replicated and EC pools on one clusterReplication vs Erasure Coding · advanced · ~18 min

Part XXVIII

Ceph Networking

Public network, cluster network where relevant, client, replication, recovery, heartbeat, MTU, bonding, VLANs.

6 lessons
  1. 01The public network: what crosses itCeph Networking · intermediate · ~15 min
  2. 02The cluster network: when to have oneCeph Networking · advanced · ~17 min
  3. 03Client traffic characteristics by interfaceCeph Networking · intermediate · ~16 min
  4. 04Replication traffic and what it multipliesCeph Networking · intermediate · ~16 min
  5. 05Recovery traffic and the throttles that shape itCeph Networking · advanced · ~18 min
  6. 06MTU, bonding, and VLANs: the link layer under CephCeph Networking · advanced · ~18 min

Part XXIX

Network Design

10/25/40/100 GbE, redundancy, switch failure domains, LACP, routed designs.

6 lessons
  1. 01Choosing link speed for a Ceph clusterNetwork Design · intermediate · ~16 min
  2. 02Bonding for redundancy and capacityNetwork Design · intermediate · ~16 min
  3. 03Switches as failure domainsNetwork Design · advanced · ~17 min
  4. 04LACP in detail: modes, rate, and hashingNetwork Design · advanced · ~16 min
  5. 05Routed network designs for CephNetwork Design · expert · ~17 min
  6. 06Jumbo frames: benefit, risk, and verificationNetwork Design · intermediate · ~16 min

Part XXX

Network Failure Behaviour

Packet loss, latency, asymmetric paths, switch failure, bond member failure. Storage symptoms.

6 lessons
  1. 01Packet loss and why it destabilises OSDsNetwork Failure Behaviour · advanced · ~17 min
  2. 02How network latency propagates into storage latencyNetwork Failure Behaviour · intermediate · ~16 min
  3. 03Asymmetric paths and one-way problemsNetwork Failure Behaviour · expert · ~17 min
  4. 04Switch failure: what Ceph seesNetwork Failure Behaviour · advanced · ~17 min
  5. 05Bond member failure versus bond failureNetwork Failure Behaviour · intermediate · ~15 min
  6. 06MTU mismatch: the most baffling slow-ops causeNetwork Failure Behaviour · advanced · ~16 min

Part XXXI

Ceph Authentication (cephx)

cephx, entities, keyrings, capabilities.

6 lessons
  1. 01The cephx authentication protocolCeph Authentication (cephx) · intermediate · ~17 min
  2. 02Entities: users and daemons in the cephx namespaceCeph Authentication (cephx) · intermediate · ~15 min
  3. 03Keyrings: format, locations, and handlingCeph Authentication (cephx) · intermediate · ~15 min
  4. 04Capabilities: what an entity is permitted to doCeph Authentication (cephx) · intermediate · ~16 min
  5. 05Capability semantics: r, w, x, and scopingCeph Authentication (cephx) · advanced · ~17 min
  6. 06Rotating cephx keys without an outageCeph Authentication (cephx) · advanced · ~18 min

Part XXXII

Least Privilege Capabilities

Capability design for RBD clients, CephFS, RGW, monitoring, administration.

6 lessons
  1. 01Minimal capabilities for RBD clientsLeast Privilege Capabilities · intermediate · ~16 min
  2. 02Minimal capabilities for CephFS clientsLeast Privilege Capabilities · advanced · ~17 min
  3. 03RGW identity: two separate systemsLeast Privilege Capabilities · intermediate · ~17 min
  4. 04Capabilities for monitoring and metrics clientsLeast Privilege Capabilities · intermediate · ~15 min
  5. 05client.admin and the discipline around itLeast Privilege Capabilities · advanced · ~16 min
  6. 06Running a capability auditLeast Privilege Capabilities · advanced · ~18 min

Part XXXIII

Encryption

Network encryption where supported, disk and device encryption, application-level encryption, RGW SSE.

6 lessons
  1. 01msgr2 and encryption on the wireEncryption · advanced · ~17 min
  2. 02Encryption at rest with dm-crypt and self-encrypting drivesEncryption · advanced · ~18 min
  3. 03Deploying secure mode across a clusterEncryption · advanced · ~17 min
  4. 04crc mode: what it is for and its limitsEncryption · intermediate · ~15 min
  5. 05RGW server-side encryption: SSE-S3, SSE-KMS, and SSE-CEncryption · advanced · ~18 min
  6. 06Protecting monitor state at restEncryption · advanced · ~17 min

Part XXXIV

Multi-Tenancy Concepts

Isolation across RBD pools, CephFS subvolumes, RGW users and tenants.

6 lessons
  1. 01Tenant isolation for RBD: pools and namespacesMulti-Tenancy Concepts · advanced · ~17 min
  2. 02CephFS subvolumes and multi-tenant file storageMulti-Tenancy Concepts · advanced · ~17 min
  3. 03RGW tenants and bucket namespace separationMulti-Tenancy Concepts · advanced · ~17 min
  4. 04RGW quotas: user, bucket, and defaultsMulti-Tenancy Concepts · intermediate · ~16 min
  5. 05Matching the isolation mechanism to the tenantMulti-Tenancy Concepts · advanced · ~17 min
  6. 06Testing tenant isolation before launchMulti-Tenancy Concepts · advanced · ~18 min

Part XXXV

RBD Architecture

VM / host to RBD client to RADOS to pool to PG to OSDs.

6 lessons
  1. 01The RBD I/O path from guest to OSDRBD Architecture · intermediate · ~17 min
  2. 02Image layout: objects, striping, and sparsenessRBD Architecture · advanced · ~17 min
  3. 03RBD image features and what each costsRBD Architecture · advanced · ~17 min
  4. 04RBD snapshots and copy-on-writeRBD Architecture · intermediate · ~17 min
  5. 05Clones: copy-on-write images from snapshotsRBD Architecture · advanced · ~17 min
  6. 06Three ways to attach an RBD imageRBD Architecture · advanced · ~18 min

Part XXXVI

RBD Images

Images, features, snapshots, clones, flattening, resizing.

6 lessons
  1. 01Creating RBD images with the right parametersRBD Images · foundation · ~15 min
  2. 02Choosing a feature set per image classRBD Images · intermediate · ~16 min
  3. 03Flattening clones and managing the dependency graphRBD Images · advanced · ~17 min
  4. 04Resizing RBD images, up and downRBD Images · intermediate · ~16 min
  5. 05Image metadata and the objects behind itRBD Images · advanced · ~17 min
  6. 06Snapshots, clones, and images: what each one isRBD Images · intermediate · ~16 min

Part XXXVII

RBD Snapshots

Snapshots are not backup. Same-cluster dependency, failure domains, rollback implications.

6 lessons
  1. 01Taking snapshots that are actually usableRBD Snapshots · intermediate · ~16 min
  2. 02Rolling back to a snapshotRBD Snapshots · intermediate · ~16 min
  3. 03What snapshots protect against, and what they do notRBD Snapshots · intermediate · ~16 min
  4. 04Snapshot protection and its role in clone safetyRBD Snapshots · intermediate · ~15 min
  5. 05The snapshot and clone lifecycle end to endRBD Snapshots · advanced · ~18 min
  6. 06The performance cost of snapshots and clone depthRBD Snapshots · advanced · ~17 min

Part XXXVIII

RBD Performance

Queue depth, caching, workload block size, replication overhead, network, OSD latency.

6 lessons
  1. 01Queue depth: the parameter that decides RBD throughputRBD Performance · advanced · ~17 min
  2. 02librbd caching and what it is safe to enableRBD Performance · advanced · ~17 min
  3. 03Block size and the layers that have oneRBD Performance · advanced · ~17 min
  4. 04What replication adds to every RBD writeRBD Performance · intermediate · ~16 min
  5. 05Why one slow OSD dominates RBD latencyRBD Performance · advanced · ~17 min
  6. 06The network component of RBD latencyRBD Performance · intermediate · ~16 min

Part XXXIX

RBD Troubleshooting

Image unavailable, client auth, mapping, latency, blocked requests, client and kernel version compatibility.

6 lessons
  1. 01When an image will not openRBD Troubleshooting · advanced · ~17 min
  2. 02Diagnosing RBD authentication failuresRBD Troubleshooting · intermediate · ~16 min
  3. 03When rbd map failsRBD Troubleshooting · advanced · ~17 min
  4. 04Working an RBD latency complaint to its causeRBD Troubleshooting · advanced · ~18 min
  5. 05Blocked requests and full-cluster conditionsRBD Troubleshooting · advanced · ~18 min
  6. 06Kernel client and librbd: differences that matter in an incidentRBD Troubleshooting · advanced · ~17 min

Part XL

CephFS Architecture

Client to MDS to metadata to RADOS data pools. Metadata vs file contents.

6 lessons
  1. 01The CephFS split path: metadata through the MDS, data directCephFS Architecture · intermediate · ~17 min
  2. 02The metadata pool and the data poolsCephFS Architecture · advanced · ~17 min
  3. 03File layouts: how a CephFS file maps to objectsCephFS Architecture · advanced · ~17 min
  4. 04MDS ranks and how the namespace is dividedCephFS Architecture · advanced · ~17 min
  5. 05The MDS journal and why recovery worksCephFS Architecture · expert · ~18 min
  6. 06CephFS capabilities: the client caching contractCephFS Architecture · expert · ~18 min

Part XLI

Metadata Servers (MDS)

Active, standby, metadata cache, failover. MDS rank 0, recovery, session behaviour.

6 lessons
  1. 01Active and standby MDS daemonsMetadata Servers (MDS) · intermediate · ~16 min
  2. 02How ranks divide one namespaceMetadata Servers (MDS) · expert · ~18 min
  3. 03The MDS cache and its memory limitMetadata Servers (MDS) · advanced · ~17 min
  4. 04MDS failover: the sequence and its timingsMetadata Servers (MDS) · advanced · ~17 min
  5. 05What happens during MDS recoveryMetadata Servers (MDS) · expert · ~18 min
  6. 06Why CephFS does not have MDS split brainMetadata Servers (MDS) · expert · ~18 min

Part XLII

CephFS Operations

Filesystem creation, clients, quotas, snapshots, subvolumes.

6 lessons
  1. 01Creating a CephFS filesystem correctly the first timeCephFS Operations · intermediate · ~16 min
  2. 02Mounting CephFS: kernel and FUSECephFS Operations · intermediate · ~17 min
  3. 03CephFS quotas and their enforcement modelCephFS Operations · intermediate · ~16 min
  4. 04CephFS snapshots and the .snap directoryCephFS Operations · advanced · ~17 min
  5. 05Subvolumes as the operational unitCephFS Operations · advanced · ~17 min
  6. 06CephFS client behaviour and consistencyCephFS Operations · advanced · ~18 min

Part XLIII

CephFS Failure Scenarios

MDS failure, slow metadata operations, full pools, client issues, session timeouts.

6 lessons
  1. 01Responding to an MDS daemon crashCephFS Failure Scenarios · advanced · ~17 min
  2. 02Diagnosing a slow MDSCephFS Failure Scenarios · advanced · ~18 min
  3. 03When a CephFS pool fillsCephFS Failure Scenarios · advanced · ~17 min
  4. 04When a client sees the wrong thingCephFS Failure Scenarios · advanced · ~17 min
  5. 05MDS journal damage and recoveryCephFS Failure Scenarios · expert · ~20 min
  6. 06Session timeouts, eviction, and stale file handlesCephFS Failure Scenarios · advanced · ~17 min

Part XLIV

Object Storage Foundations

Objects, buckets, S3 concepts, object namespaces, versioning, lifecycle.

6 lessons
  1. 01What object storage actually promisesObject Storage Foundations · foundation · ~16 min
  2. 02Buckets: the unit of policy and the unit of scaleObject Storage Foundations · intermediate · ~17 min
  3. 03The S3 API surface RGW implementsObject Storage Foundations · intermediate · ~17 min
  4. 04Flat namespaces and the prefix conventionObject Storage Foundations · intermediate · ~16 min
  5. 05S3 versioning and its capacity consequencesObject Storage Foundations · advanced · ~17 min
  6. 06Lifecycle rules: automated expiration and transitionObject Storage Foundations · advanced · ~18 min

Part XLV

RADOS Gateway (RGW)

S3 client to RGW to RADOS architecture, daemon placement, multisite concepts.

6 lessons
  1. 01RGW: translating S3 into RADOSRADOS Gateway (RGW) · intermediate · ~17 min
  2. 02Placing and sizing RGW daemonsRADOS Gateway (RGW) · intermediate · ~16 min
  3. 03Zones: the unit of data placementRADOS Gateway (RGW) · advanced · ~17 min
  4. 04Zonegroups and the shape of multisiteRADOS Gateway (RGW) · expert · ~18 min
  5. 05The RGW pools and what each requiresRADOS Gateway (RGW) · advanced · ~17 min
  6. 06High availability for RGWRADOS Gateway (RGW) · advanced · ~17 min

Part XLVI

RGW Users and Credentials

Access keys, secret keys, users, subusers, quotas, bucket policies.

6 lessons
  1. 01Creating and managing RGW usersRGW Users and Credentials · foundation · ~15 min
  2. 02Access keys, rotation, and credential handlingRGW Users and Credentials · intermediate · ~16 min
  3. 03Subusers and scoped credentialsRGW Users and Credentials · intermediate · ~16 min
  4. 04Applying RGW quotas in practiceRGW Users and Credentials · intermediate · ~16 min
  5. 05Bucket policies: the general access control mechanismRGW Users and Credentials · advanced · ~18 min
  6. 06RGW admin capabilitiesRGW Users and Credentials · advanced · ~17 min

Part XLVII

RGW High Availability

Multiple gateways, load balancing, zonegroup / zone placement, eventual consistency in multisite.

6 lessons
  1. 01Running multiple gateways wellRGW High Availability · intermediate · ~16 min
  2. 02Load balancing RGW: layer 4 versus layer 7RGW High Availability · advanced · ~17 min
  3. 03Setting up RGW multisiteRGW High Availability · expert · ~19 min
  4. 04Sync policies: controlling what replicatesRGW High Availability · expert · ~19 min
  5. 05Eventual consistency across zonesRGW High Availability · expert · ~18 min
  6. 06RGW failure modes and their responsesRGW High Availability · advanced · ~18 min

Part XLVIII

RGW Troubleshooting

Authentication, bucket access, backend pool, latency, load balancer.

6 lessons
  1. 01S3 authentication failures and what each meansRGW Troubleshooting · intermediate · ~17 min
  2. 02Working an AccessDenied to its sourceRGW Troubleshooting · advanced · ~17 min
  3. 03When the pools behind RGW are the problemRGW Troubleshooting · advanced · ~17 min
  4. 04Tracing RGW latency to its layerRGW Troubleshooting · advanced · ~18 min
  5. 05Load balancer failures in front of RGWRGW Troubleshooting · advanced · ~17 min
  6. 06Diagnosing multisite sync problemsRGW Troubleshooting · expert · ~19 min

Part XLIX

cephadm

Bootstrap, hosts, services, daemons, inventory, labels, the cephadm orchestrator.

6 lessons
  1. 01Bootstrapping a cluster with cephadmcephadm · intermediate · ~17 min
  2. 02Adding and managing hostscephadm · intermediate · ~17 min
  3. 03Declaring services with the orchestratorcephadm · intermediate · ~17 min
  4. 04Reading cluster state through the orchestratorcephadm · intermediate · ~16 min
  5. 05Device inventory and OSD creationcephadm · advanced · ~18 min
  6. 06Labels as the placement mechanismcephadm · intermediate · ~16 min

Part L

Cluster Deployment

Production deployment guidance. Node count, MON placement, MGR, OSD layout, networks, time, DNS, storage devices.

6 lessons
  1. 01How many hosts a Ceph cluster needsCluster Deployment · intermediate · ~17 min
  2. 02Placing monitorsCluster Deployment · intermediate · ~17 min
  3. 03Placing managersCluster Deployment · intermediate · ~15 min
  4. 04OSD layout: devices, hosts, and what to co-locateCluster Deployment · advanced · ~18 min
  5. 05Network design for a production clusterCluster Deployment · advanced · ~17 min
  6. 06Time and name resolution as deployment prerequisitesCluster Deployment · intermediate · ~16 min

Part LI

Host Preparation

Linux, networking, time sync, hostname / DNS, disks, container runtime, partitions, device-by-id.

6 lessons
  1. 01The Linux baseline for a Ceph hostHost Preparation · intermediate · ~17 min
  2. 02Host networking configuration for CephHost Preparation · intermediate · ~17 min
  3. 03Configuring time synchronisation for CephHost Preparation · intermediate · ~15 min
  4. 04Hostnames and resolution on Ceph hostsHost Preparation · intermediate · ~15 min
  5. 05Preparing disks for OSD useHost Preparation · intermediate · ~17 min
  6. 06The container runtime under cephadmHost Preparation · advanced · ~17 min

Part LII

Time Synchronisation

Time as a production dependency. Consequences for distributed systems, MON election, recovery, log correlation.

6 lessons
  1. 01What breaks in Ceph when time is wrongTime Synchronisation · intermediate · ~16 min
  2. 02Monitor elections and clock skewTime Synchronisation · advanced · ~18 min
  3. 03Certificates, msgr2, and timeTime Synchronisation · intermediate · ~16 min
  4. 04Log correlation and why it needs synchronised timeTime Synchronisation · intermediate · ~16 min
  5. 05Choosing and configuring a time sourceTime Synchronisation · intermediate · ~16 min
  6. 06A time skew incident, worked throughTime Synchronisation · advanced · ~18 min

Part LIII

Cluster Health

HEALTH_OK, HEALTH_WARN, HEALTH_ERR. Reading health messages, not just health states.

6 lessons
  1. 01The three health states and what they actually meanCluster Health · foundation · ~15 min
  2. 02What HEALTH_OK does not tell youCluster Health · intermediate · ~16 min
  3. 03Reading HEALTH_WARN correctlyCluster Health · intermediate · ~17 min
  4. 04HEALTH_ERR and the checks that produce itCluster Health · advanced · ~18 min
  5. 05Getting the most from ceph health detailCluster Health · intermediate · ~16 min
  6. 06A reference for the health checks you will actually seeCluster Health · intermediate · ~18 min

Part LIV

ceph status and health detail

ceph -s, ceph health detail, ceph health with output interpretation.

6 lessons
  1. 01Reading ceph -s line by lineceph status and health detail · foundation · ~16 min
  2. 02ceph health, detail, and the mute mechanismceph status and health detail · intermediate · ~15 min
  3. 03Inspecting monitor quorumceph status and health detail · advanced · ~16 min
  4. 04Monitor status at a glanceceph status and health detail · intermediate · ~15 min
  5. 05Manager status and modulesceph status and health detail · intermediate · ~15 min
  6. 06PG statistics as a progress meterceph status and health detail · intermediate · ~15 min

Part LV

OSD States

up / down and in / out. Why these are separate dimensions. Critical lesson.

6 lessons
  1. 01up and down: the liveness axisOSD States · foundation · ~16 min
  2. 02in and out: the placement axisOSD States · foundation · ~16 min
  3. 03The four OSD states and what each meansOSD States · intermediate · ~16 min
  4. 04Why liveness and placement are separate concernsOSD States · intermediate · ~16 min
  5. 05The lifecycle of an OSD through its statesOSD States · intermediate · ~17 min
  6. 06The safe order for OSD operationsOSD States · advanced · ~17 min

Part LVI

OSD Failure

OSD down to PG degraded to recovery decisions to OSD out to backfill. The full failure cascade.

6 lessons
  1. 01The first minutes of an OSD failureOSD Failure · intermediate · ~16 min
  2. 02What degraded PGs mean after an OSD failureOSD Failure · intermediate · ~16 min
  3. 03Deciding what to do about a failed OSDOSD Failure · advanced · ~17 min
  4. 04Marking an OSD out and watching the rebalanceOSD Failure · intermediate · ~16 min
  5. 05Stopping the OSD daemon at the right momentOSD Failure · intermediate · ~15 min
  6. 06Purging an OSD from the clusterOSD Failure · intermediate · ~16 min

Part LVII

Replacing Failed OSDs

Identify disk, verify failure, safely remove, replace, recreate, monitor recovery.

6 lessons
  1. 01Identifying the physical disk behind a failed OSDReplacing Failed OSDs · intermediate · ~17 min
  2. 02Draining before replacingReplacing Failed OSDs · intermediate · ~16 min
  3. 03Stopping and purging before the swapReplacing Failed OSDs · intermediate · ~15 min
  4. 04The physical replacementReplacing Failed OSDs · intermediate · ~16 min
  5. 05Creating the replacement OSDReplacing Failed OSDs · advanced · ~17 min
  6. 06Watching the new OSD fillReplacing Failed OSDs · intermediate · ~16 min

Part LVIII

Recovery

Recovery semantics, missing replicas, object reconstruction, recovery priorities, impact on production I/O.

6 lessons
  1. 01What recovery does and how it knows what to copyRecovery · intermediate · ~17 min
  2. 02The settings that bound recovery speedRecovery · advanced · ~17 min
  3. 03How recovery affects client I/ORecovery · advanced · ~17 min
  4. 04Per-pool recovery priorityRecovery · advanced · ~16 min
  5. 05Backfill priority and how it differs from recovery priorityRecovery · advanced · ~16 min
  6. 06Knowing when recovery is finishedRecovery · intermediate · ~15 min

Part LIX

Backfill

Why topology and capacity changes trigger data movement. Difference between recovery and backfill.

6 lessons
  1. 01Backfill and recovery: two mechanisms, two triggersBackfill · intermediate · ~16 min
  2. 02Every change that triggers backfillBackfill · advanced · ~17 min
  3. 03Controlling backfill speedBackfill · advanced · ~17 min
  4. 04The client impact of backfill and how to bound itBackfill · advanced · ~17 min
  5. 05Pausing and resuming backfillBackfill · intermediate · ~16 min
  6. 06Estimating how long a backfill will takeBackfill · intermediate · ~16 min

Part LX

Recovery Tuning

Balancing client I/O vs recovery speed. osd_recovery_max_active, osd_recovery_sleep, osd_max_backfills.

6 lessons
  1. 01The client versus recovery trade-offRecovery Tuning · intermediate · ~16 min
  2. 02The recovery settings and what each doesRecovery Tuning · advanced · ~17 min
  3. 03The backfill settings and their interactionRecovery Tuning · advanced · ~17 min
  4. 04Per-pool recovery priorityRecovery Tuning · intermediate · ~16 min
  5. 05The balancer and upmapRecovery Tuning · advanced · ~18 min
  6. 06Running a recovery tuning cycleRecovery Tuning · advanced · ~17 min

Part LXI

Scrubbing

Scrub vs deep scrub. Checksums, consistency, daily and weekly schedules, scrubbing performance impact.

6 lessons
  1. 01Scrub and deep scrub: what each actually checksScrubbing · intermediate · ~16 min
  2. 02How Ceph schedules scrubsScrubbing · intermediate · ~17 min
  3. 03Scrub impact and how to shape itScrubbing · advanced · ~17 min
  4. 04Tuning scrub deliberatelyScrubbing · advanced · ~18 min
  5. 05When a scrub finds a problemScrubbing · advanced · ~17 min
  6. 06Pausing scrubScrubbing · intermediate · ~15 min

Part LXII

Inconsistent PGs

How inconsistent objects are detected and repaired. Strong warnings around pg repair.

6 lessons
  1. 01Detecting an inconsistent PGInconsistent PGs · intermediate · ~16 min
  2. 02Diagnosing the inconsistencyInconsistent PGs · advanced · ~18 min
  3. 03Deciding how to repairInconsistent PGs · advanced · ~17 min
  4. 04Running the repairInconsistent PGs · intermediate · ~16 min
  5. 05Verifying a repairInconsistent PGs · intermediate · ~16 min
  6. 06Preventing inconsistenciesInconsistent PGs · intermediate · ~17 min

Part LXIII

Capacity Management

Raw, usable, replication overhead, EC overhead, reserved headroom, recovery capacity, backfill capacity.

6 lessons
  1. 01Raw capacity and usable capacityCapacity Management · intermediate · ~17 min
  2. 02The cost of replicationCapacity Management · intermediate · ~16 min
  3. 03The cost of erasure codingCapacity Management · advanced · ~17 min
  4. 04How much free space to keepCapacity Management · intermediate · ~17 min
  5. 05Capacity for recovery and backfillCapacity Management · advanced · ~17 min
  6. 06Monitoring capacityCapacity Management · intermediate · ~17 min

Part LXIV

Nearfull, Backfillfull and Full

Current threshold behaviour. Why clusters need free capacity to recover.

6 lessons
  1. 01The nearfull thresholdNearfull, Backfillfull and Full · intermediate · ~16 min
  2. 02The backfillfull thresholdNearfull, Backfillfull and Full · advanced · ~17 min
  3. 03The full thresholdNearfull, Backfillfull and Full · advanced · ~17 min
  4. 04Why a full cluster cannot heal itselfNearfull, Backfillfull and Full · advanced · ~17 min
  5. 05Getting out of nearfullNearfull, Backfillfull and Full · advanced · ~18 min
  6. 06Alerting on capacity thresholdsNearfull, Backfillfull and Full · intermediate · ~17 min

Part LXV

Why Full Clusters Are Dangerous

Cluster becomes full, writes blocked, recovery constrained, operational options shrink.

6 lessons
  1. 01What blocked writes look like from the applicationWhy Full Clusters Are Dangerous · intermediate · ~17 min
  2. 02Recovery under capacity constraintWhy Full Clusters Are Dangerous · advanced · ~17 min
  3. 03How a full cluster narrows your optionsWhy Full Clusters Are Dangerous · intermediate · ~16 min
  4. 04Emergency space reclamationWhy Full Clusters Are Dangerous · advanced · ~18 min
  5. 05The full-cluster trap and how it formsWhy Full Clusters Are Dangerous · advanced · ~17 min
  6. 06Recovering a cluster after a full eventWhy Full Clusters Are Dangerous · advanced · ~18 min

Part LXVI

Capacity Forecasting

Growth rate, retention, workload expansion, recovery headroom, node-addition lead time.

6 lessons
  1. 01Measuring growth and computing lead timeCapacity Forecasting · intermediate · ~17 min
  2. 02Retention and churnCapacity Forecasting · intermediate · ~17 min
  3. 03Planning for new workloadsCapacity Forecasting · intermediate · ~17 min
  4. 04Headroom for a host failureCapacity Forecasting · advanced · ~17 min
  5. 05Headroom for maintenanceCapacity Forecasting · intermediate · ~16 min
  6. 06Keeping a forecast accurateCapacity Forecasting · intermediate · ~17 min

Part LXVII

Performance Methodology

Application to client to network to OSD to device. Systematic isolation.

6 lessons
  1. 01The layered model for performance workPerformance Methodology · intermediate · ~17 min
  2. 02Establishing a performance baselinePerformance Methodology · intermediate · ~17 min
  3. 03Isolating the slow layerPerformance Methodology · advanced · ~18 min
  4. 04Correlating evidence across layersPerformance Methodology · advanced · ~18 min
  5. 05Working a slow-ops reportPerformance Methodology · advanced · ~18 min
  6. 06Working a blocked-ops reportPerformance Methodology · advanced · ~18 min

Part LXVIII

OSD Latency

Identifying slow OSDs, correlating ceph daemonperf, perf dump, OSD latency histograms.

6 lessons
  1. 01Finding the slow OSDOSD Latency · intermediate · ~17 min
  2. 02Connecting OSD latency to what users seeOSD Latency · advanced · ~17 min
  3. 03Tuning OSD latencyOSD Latency · advanced · ~18 min
  4. 04The device latency floorOSD Latency · intermediate · ~17 min
  5. 05Contention between workloads on shared OSDsOSD Latency · advanced · ~18 min
  6. 06Bounding recovery's cost to client latencyOSD Latency · advanced · ~17 min

Part LXIX

Disk Performance

Disk-side latency, queueing, device errors, SMART / NVMe health integration with Ceph.

6 lessons
  1. 01Choosing a device class for a workloadDisk Performance · intermediate · ~17 min
  2. 02Queue depth and where latency accumulatesDisk Performance · advanced · ~17 min
  3. 03Reading device health dataDisk Performance · intermediate · ~18 min
  4. 04Flash wear and write amplificationDisk Performance · advanced · ~18 min
  5. 05Automating device health collectionDisk Performance · intermediate · ~17 min
  6. 06Deciding when to replace a deviceDisk Performance · intermediate · ~17 min

Part LXX

Network Performance

Storage latency can actually be network latency. Bonding, MTU, asymmetric routes, switch load.

6 lessons
  1. 01When storage latency is really network latencyNetwork Performance · advanced · ~17 min
  2. 02MTU end to endNetwork Performance · intermediate · ~17 min
  3. 03Bonded links and what happens when one failsNetwork Performance · advanced · ~17 min
  4. 04Switch saturation and bufferingNetwork Performance · advanced · ~18 min
  5. 05Asymmetric and unequal pathsNetwork Performance · advanced · ~17 min
  6. 06Ceph's network traffic patternsNetwork Performance · intermediate · ~17 min

Part LXXI

Client Performance

RBD client behaviour, queue depth, concurrency, caching, krbd vs librbd vs rbd-nbd.

6 lessons
  1. 01librbd client configurationClient Performance · intermediate · ~17 min
  2. 02Kernel RBD and librbd comparedClient Performance · intermediate · ~17 min
  3. 03rbd-nbd and when it is the right choiceClient Performance · intermediate · ~16 min
  4. 04Concurrency and where throughput comes fromClient Performance · advanced · ~17 min
  5. 05Choosing a cache modeClient Performance · advanced · ~17 min
  6. 06Client-side queueing and threadingClient Performance · advanced · ~17 min

Part LXXII

Benchmarking

rados bench, rbd bench, fio against RBD, ceph bench where appropriate. Never against production.

6 lessons
  1. 01rados benchBenchmarking · intermediate · ~17 min
  2. 02rbd benchBenchmarking · intermediate · ~16 min
  3. 03fio against RBDBenchmarking · advanced · ~18 min
  4. 04The per-OSD benchmarkBenchmarking · advanced · ~17 min
  5. 05Benchmarking without disturbing productionBenchmarking · intermediate · ~17 min
  6. 06What to capture from a benchmarkBenchmarking · intermediate · ~17 min

Part LXXIII

Benchmark Interpretation

Latency distribution, IOPS, throughput, queue depth, unrealistic benchmarks, workload mismatch.

6 lessons
  1. 01Reading a latency distributionBenchmark Interpretation · advanced · ~17 min
  2. 02IOPS and block sizeBenchmark Interpretation · intermediate · ~16 min
  3. 03Throughput and saturationBenchmark Interpretation · advanced · ~17 min
  4. 04Benchmarks that misleadBenchmark Interpretation · advanced · ~18 min
  5. 05Benchmarking the workload you actually haveBenchmark Interpretation · advanced · ~18 min
  6. 06Tracing tail latency to its sourceBenchmark Interpretation · advanced · ~18 min

Part LXXIV

Observability

Monitoring cluster health, MON quorum, OSD states, PG states, capacity, latency, recovery, network, device health.

6 lessons
  1. 01Getting Ceph metrics into PrometheusObservability · intermediate · ~17 min
  2. 02The metrics worth watchingObservability · intermediate · ~18 min
  3. 03Health status as a metricObservability · intermediate · ~17 min
  4. 04Monitoring the monitorsObservability · advanced · ~17 min
  5. 05OSD state metricsObservability · intermediate · ~17 min
  6. 06PG state metricsObservability · advanced · ~17 min

Part LXXV

Prometheus Metrics

MGR Prometheus exporter, high-value metrics, histograms vs counters, latency observation.

6 lessons
  1. 01Operating the metrics exporterPrometheus Metrics · advanced · ~17 min
  2. 02Building useful queries from raw metricsPrometheus Metrics · advanced · ~18 min
  3. 03Labels and cardinalityPrometheus Metrics · advanced · ~17 min
  4. 04Recording rules for CephPrometheus Metrics · advanced · ~17 min
  5. 05Writing Ceph alerting rulesPrometheus Metrics · advanced · ~18 min
  6. 06Scrape configuration and reliabilityPrometheus Metrics · intermediate · ~17 min

Part LXXVI

Grafana Dashboards

Per-cluster health, per-pool IOPS, per-OSD latency, recovery, capacity, network.

6 lessons
  1. 01The cluster overview dashboardGrafana Dashboards · intermediate · ~17 min
  2. 02The per-pool dashboardGrafana Dashboards · intermediate · ~17 min
  3. 03The per-OSD dashboardGrafana Dashboards · advanced · ~17 min
  4. 04The recovery dashboardGrafana Dashboards · intermediate · ~17 min
  5. 05The capacity dashboardGrafana Dashboards · intermediate · ~17 min
  6. 06The host and network dashboardGrafana Dashboards · advanced · ~17 min

Part LXXVII

Alerting

Actionable alerts: MON quorum loss, OSD down, degraded PGs, inconsistent PG, nearfull, slow ops.

6 lessons
  1. 01The monitor quorum alert and its responseAlerting · advanced · ~17 min
  2. 02The OSD down alert and its responseAlerting · intermediate · ~17 min
  3. 03Degraded PG alerts and their triageAlerting · advanced · ~17 min
  4. 04The inconsistent PG alert and its handlingAlerting · advanced · ~17 min
  5. 05The nearfull alert and the capacity conversationAlerting · intermediate · ~17 min
  6. 06The slow ops alert and its escalationAlerting · advanced · ~17 min

Part LXXVIII

Monitoring the Monitoring Path

What happens if Ceph is failing and the telemetry pipeline is also unavailable.

6 lessons
  1. 01Every part of the telemetry path can fail silentlyMonitoring the Monitoring Path · advanced · ~17 min
  2. 02Redundant Prometheus for Ceph monitoringMonitoring the Monitoring Path · advanced · ~17 min
  3. 03Redundant AlertmanagerMonitoring the Monitoring Path · advanced · ~17 min
  4. 04Grafana availability and what it actually needsMonitoring the Monitoring Path · intermediate · ~16 min
  5. 05Scaling the monitoring stack with the clusterMonitoring the Monitoring Path · advanced · ~17 min
  6. 06Operating the monitoring stackMonitoring the Monitoring Path · intermediate · ~17 min

Part LXXIX

Slow Ops

What slow ops mean. OSD, client, network causes. Correlating evidence across slow ops and daemonperf.

6 lessons
  1. 01What Ceph counts as a slow operationSlow Ops · intermediate · ~16 min
  2. 02Slow ops originating at an OSDSlow Ops · advanced · ~18 min
  3. 03Slow ops originating at the clientSlow Ops · advanced · ~17 min
  4. 04Slow ops originating in the networkSlow Ops · advanced · ~17 min
  5. 05Slow ops caused by recoverySlow Ops · advanced · ~17 min
  6. 06Correlating slow ops with cluster stateSlow Ops · advanced · ~18 min

Part LXXX

Blocked Operations

blocked vs slow ops. Diagnostic interpretation.

6 lessons
  1. 01Blocked and slow: the distinction that changes everythingBlocked Operations · intermediate · ~16 min
  2. 02Blocked by capacity thresholdsBlocked Operations · advanced · ~17 min
  3. 03Pool-level blocking conditionsBlocked Operations · intermediate · ~17 min
  4. 04Blocked by insufficient available copiesBlocked Operations · advanced · ~18 min
  5. 05Blocked by forgotten flagsBlocked Operations · intermediate · ~17 min
  6. 06The blocked operations sweepBlocked Operations · advanced · ~17 min

Part LXXXI

Proxmox Integration

RBD storage, CephFS where applicable, hyper-converged design, dedicated storage clusters, networking.

6 lessons
  1. 01RBD as Proxmox VE storageProxmox Integration · intermediate · ~17 min
  2. 02CephFS as Proxmox storageProxmox Integration · intermediate · ~17 min
  3. 03Hyper-converged Proxmox and CephProxmox Integration · advanced · ~18 min
  4. 04A dedicated Ceph cluster for ProxmoxProxmox Integration · intermediate · ~17 min
  5. 05Network design for Proxmox with CephProxmox Integration · advanced · ~17 min
  6. 06Failure domains for Proxmox workloadsProxmox Integration · advanced · ~18 min

Part LXXXII

Hyper-Converged Ceph

Compute plus Ceph OSD on same nodes. Cost, performance contention, failure domains, maintenance.

6 lessons
  1. 01The economics of hyper-convergenceHyper-Converged Ceph · intermediate · ~17 min
  2. 02Managing compute and storage contentionHyper-Converged Ceph · advanced · ~18 min
  3. 03Correlated failure in a hyper-converged clusterHyper-Converged Ceph · advanced · ~18 min
  4. 04Maintenance on a hyper-converged nodeHyper-Converged Ceph · advanced · ~18 min
  5. 05Scaling a hyper-converged clusterHyper-Converged Ceph · intermediate · ~17 min
  6. 06Migrating from hyper-converged to dedicatedHyper-Converged Ceph · advanced · ~18 min

Part LXXXIII

Dedicated Ceph Cluster

Dedicated storage nodes vs hyper-converged. Trade-offs.

6 lessons
  1. 01Designing a dedicated storage clusterDedicated Ceph Cluster · advanced · ~18 min
  2. 02Network design for a dedicated clusterDedicated Ceph Cluster · advanced · ~17 min
  3. 03Monitoring a dedicated clusterDedicated Ceph Cluster · intermediate · ~17 min
  4. 04Running a dedicated cluster as a serviceDedicated Ceph Cluster · intermediate · ~18 min
  5. 05Comparing the architectures honestlyDedicated Ceph Cluster · intermediate · ~17 min
  6. 06Serving multiple consumers from one clusterDedicated Ceph Cluster · advanced · ~18 min

Part LXXXIV

Proxmox Failure Scenarios

VM I/O latency, failed OSD, node maintenance, quorum dependencies.

6 lessons
  1. 01Diagnosing VM disk latency on CephProxmox Failure Scenarios · advanced · ~18 min
  2. 02An OSD failure and its effect on running VMsProxmox Failure Scenarios · advanced · ~17 min
  3. 03Coordinating Proxmox and Ceph maintenanceProxmox Failure Scenarios · advanced · ~17 min
  4. 04Ceph quorum loss with VMs runningProxmox Failure Scenarios · advanced · ~17 min
  5. 05Path redundancy with RBDProxmox Failure Scenarios · intermediate · ~17 min
  6. 06Total storage loss with VMs runningProxmox Failure Scenarios · advanced · ~18 min

Part LXXXV

Kubernetes Integration

CSI (csi-rbd, csi-cephfs). StorageClass, PVC, secrets, provisioner, snapshotter, topology awareness.

6 lessons
  1. 01The CSI architecture for CephKubernetes Integration · intermediate · ~18 min
  2. 02Deploying and configuring Ceph-CSIKubernetes Integration · advanced · ~18 min
  3. 03Designing StorageClasses for CephKubernetes Integration · advanced · ~18 min
  4. 04PersistentVolumeClaims and their lifecycleKubernetes Integration · advanced · ~18 min
  5. 05Topology-aware volume placementKubernetes Integration · advanced · ~18 min
  6. 06Volume encryption with Ceph-CSIKubernetes Integration · advanced · ~18 min

Part LXXXVI

Kubernetes RBD

Block volumes via csi-rbd. StorageClass parameters, RBD image format, topology, dynamic provisioning.

6 lessons
  1. 01RBD volumes and access modesKubernetes RBD · intermediate · ~17 min
  2. 02RBD StorageClass parameters in depthKubernetes RBD · advanced · ~18 min
  3. 03RBD image layout for Kubernetes volumesKubernetes RBD · advanced · ~17 min
  4. 04RBD image features and kernel compatibilityKubernetes RBD · advanced · ~17 min
  5. 05Aligning Kubernetes and Ceph topologyKubernetes RBD · advanced · ~17 min
  6. 06CSI snapshots and what they are notKubernetes RBD · advanced · ~18 min

Part LXXXVII

Kubernetes CephFS

Shared filesystem via csi-cephfs. Subvolume provisioning, subvolume snapshots, dynamic provisioning.

6 lessons
  1. 01CephFS subvolumes as Kubernetes volumesKubernetes CephFS · intermediate · ~17 min
  2. 02CephFS snapshots through CSIKubernetes CephFS · advanced · ~17 min
  3. 03Shared CephFS volumes across podsKubernetes CephFS · advanced · ~18 min
  4. 04Permissions and identity on CephFS volumesKubernetes CephFS · advanced · ~18 min
  5. 05CephFS StorageClass configurationKubernetes CephFS · advanced · ~17 min
  6. 06CephFS failure behaviour in KubernetesKubernetes CephFS · advanced · ~18 min

Part LXXXVIII

Kubernetes Storage Failure Scenarios

PVC mount failure, CSI problems, Ceph auth, pool failure, backend capacity, stale snapshot.

6 lessons
  1. 01Diagnosing a volume that will not mountKubernetes Storage Failure Scenarios · advanced · ~18 min
  2. 02CSI driver failures and their blast radiusKubernetes Storage Failure Scenarios · advanced · ~17 min
  3. 03Authentication failures between Kubernetes and CephKubernetes Storage Failure Scenarios · advanced · ~17 min
  4. 04Kubernetes workloads during a Ceph cluster problemKubernetes Storage Failure Scenarios · advanced · ~18 min
  5. 05Stale snapshots and orphaned referencesKubernetes Storage Failure Scenarios · advanced · ~17 min
  6. 06Volume density limits per nodeKubernetes Storage Failure Scenarios · advanced · ~17 min

Part LXXXIX

Rook Concepts

Rook as a Kubernetes operator for Ceph. Forward-pointer, awareness lesson, not a deep dive.

6 lessons
  1. 01Rook: Ceph as a Kubernetes operatorRook Concepts · intermediate · ~17 min
  2. 02Rook and cephadm comparedRook Concepts · intermediate · ~17 min
  3. 03When Rook is the right choiceRook Concepts · intermediate · ~17 min
  4. 04When cephadm is the right choiceRook Concepts · intermediate · ~17 min
  5. 05Why this course teaches Ceph directlyRook Concepts · intermediate · ~16 min
  6. 06Continuing beyond this courseRook Concepts · foundation · ~16 min

Part XC

Scaling Out

Adding OSDs, hosts, capacity, MONs only when justified. Resulting rebalancing.

6 lessons
  1. 01Adding OSDs to an existing clusterScaling Out · intermediate · ~17 min
  2. 02Adding a host to the clusterScaling Out · intermediate · ~17 min
  3. 03Changing the monitor countScaling Out · advanced · ~17 min
  4. 04Manager placement and redundancyScaling Out · intermediate · ~16 min
  5. 05Managing the rebalance an expansion causesScaling Out · intermediate · ~17 min
  6. 06Expanding during degraded conditionsScaling Out · advanced · ~17 min

Part XCI

Adding Storage Nodes

Pre-checks: failure domain, network, capacity, CRUSH placement, recovery impact.

6 lessons
  1. 01Pre-check: failure domain placementAdding Storage Nodes · intermediate · ~16 min
  2. 02Pre-check: network configurationAdding Storage Nodes · intermediate · ~17 min
  3. 03Pre-check: devices and capacityAdding Storage Nodes · intermediate · ~17 min
  4. 04Pre-check: CRUSH map preparationAdding Storage Nodes · advanced · ~17 min
  5. 05Pre-check: capacity and impact for the backfillAdding Storage Nodes · intermediate · ~17 min
  6. 06Executing the node additionAdding Storage Nodes · intermediate · ~17 min

Part XCII

Removing Storage Nodes

Safe evacuation and removal. Strong warnings against simply deleting OSDs.

6 lessons
  1. 01Evacuating a nodeRemoving Storage Nodes · advanced · ~17 min
  2. 02Verifying an evacuation is completeRemoving Storage Nodes · intermediate · ~16 min
  3. 03Removing the OSDsRemoving Storage Nodes · advanced · ~17 min
  4. 04Removing the hostRemoving Storage Nodes · intermediate · ~16 min
  5. 05CRUSH map hygiene after removalsRemoving Storage Nodes · intermediate · ~16 min
  6. 06The removal order and why it existsRemoving Storage Nodes · advanced · ~17 min

Part XCIII

Changing CRUSH Topology

Blast radius, data movement impact, crush reweight, crush reweight-all, crush move.

6 lessons
  1. 01CRUSH weight and OSD reweightChanging CRUSH Topology · advanced · ~17 min
  2. 02Bulk weight correctionsChanging CRUSH Topology · advanced · ~17 min
  3. 03Moving buckets in the hierarchyChanging CRUSH Topology · advanced · ~17 min
  4. 04Changing a CRUSH ruleChanging CRUSH Topology · advanced · ~18 min
  5. 05Blast radius of a topology changeChanging CRUSH Topology · advanced · ~17 min
  6. 06Monitoring a topology changeChanging CRUSH Topology · intermediate · ~17 min

Part XCIV

Hardware Replacement

Failed disk, failing disk, failed server, replacement node.

6 lessons
  1. 01Replacing a failed diskHardware Replacement · intermediate · ~17 min
  2. 02Replacing a failing disk before it failsHardware Replacement · intermediate · ~17 min
  3. 03Replacing a failed serverHardware Replacement · advanced · ~18 min
  4. 04Bringing a replacement node into serviceHardware Replacement · intermediate · ~17 min
  5. 05Verifying replacement hardware before deploymentHardware Replacement · intermediate · ~17 min
  6. 06Validating a replacement after deploymentHardware Replacement · intermediate · ~17 min

Part XCV

Maintenance Flags

noout, norebalance, norecover, nobackfill, nodeep-scrub. Strong warning: never leave enabled indefinitely.

6 lessons
  1. 01The noout flag in depthMaintenance Flags · intermediate · ~17 min
  2. 02The norecover flag in depthMaintenance Flags · advanced · ~16 min
  3. 03The norebalance flag in depthMaintenance Flags · intermediate · ~16 min
  4. 04The nobackfill flag in depthMaintenance Flags · advanced · ~16 min
  5. 05The scrub suppression flagsMaintenance Flags · intermediate · ~16 min
  6. 06Flag lifecycle managementMaintenance Flags · intermediate · ~17 min

Part XCVI

Node Maintenance

Planned reboot, kernel update, OSD host maintenance, daemon restarts.

6 lessons
  1. 01Reboots and the down-out intervalNode Maintenance · intermediate · ~17 min
  2. 02Kernel and host software upgradesNode Maintenance · advanced · ~18 min
  3. 03What clients experience during node maintenanceNode Maintenance · intermediate · ~17 min
  4. 04Monitoring during a maintenance windowNode Maintenance · intermediate · ~16 min
  5. 05Rolling back a maintenance changeNode Maintenance · advanced · ~17 min
  6. 06Post-maintenance verificationNode Maintenance · intermediate · ~17 min

Part XCVII

Network Maintenance

Switch / link maintenance with storage redundancy, bond member replacement, MTU changes.

6 lessons
  1. 01Switch maintenance with redundant uplinksNetwork Maintenance · advanced · ~18 min
  2. 02Individual link and NIC maintenanceNetwork Maintenance · intermediate · ~16 min
  3. 03Changing MTU across a live clusterNetwork Maintenance · advanced · ~18 min
  4. 04Network maintenance during a degraded clusterNetwork Maintenance · advanced · ~17 min
  5. 05Client impact during network maintenanceNetwork Maintenance · intermediate · ~17 min
  6. 06Network change rollbackNetwork Maintenance · advanced · ~17 min

Part XCVIII

Software Upgrades

Safe Ceph upgrade workflows. Containerised cephadm upgrade path.

6 lessons
  1. 01Ceph release cadence and supported upgrade pathsSoftware Upgrades · intermediate · ~17 min
  2. 02Running a cephadm upgradeSoftware Upgrades · advanced · ~18 min
  3. 03The upgrade order and why it existsSoftware Upgrades · advanced · ~17 min
  4. 04Pre-upgrade readinessSoftware Upgrades · advanced · ~18 min
  5. 05Client version compatibilitySoftware Upgrades · advanced · ~17 min
  6. 06Upgrade rollback and the backup that replaces itSoftware Upgrades · advanced · ~18 min

Part XCIX

Upgrade Planning

Read release notes, check health, backup config/state, validate version path, upgrade staging, upgrade cluster, monitor.

6 lessons
  1. 01Reading release notes as an operational documentUpgrade Planning · intermediate · ~17 min
  2. 02Establishing the pre-upgrade baselineUpgrade Planning · intermediate · ~17 min
  3. 03What to back up before an upgrade, and what a backup cannot doUpgrade Planning · advanced · ~18 min
  4. 04Building a test cluster that predicts productionUpgrade Planning · advanced · ~18 min
  5. 05Planning and communicating the upgrade windowUpgrade Planning · intermediate · ~17 min
  6. 06The post-upgrade watch periodUpgrade Planning · intermediate · ~18 min

Part C

Upgrade Failure Recovery

Scenarios: daemon fails after upgrade, mixed-version state, module/plugin incompatibility, MGR failover.

6 lessons
  1. 01A daemon fails to start after upgradingUpgrade Failure Recovery · advanced · ~18 min
  2. 02Operating a cluster stuck in a mixed-version stateUpgrade Failure Recovery · advanced · ~18 min
  3. 03A manager module fails after upgradingUpgrade Failure Recovery · advanced · ~17 min
  4. 04Manager failover during an upgradeUpgrade Failure Recovery · intermediate · ~17 min
  5. 05A daemon crash-looping after an upgradeUpgrade Failure Recovery · advanced · ~18 min
  6. 06The post-upgrade reviewUpgrade Failure Recovery · intermediate · ~17 min

Part CI

Security Hardening

cephx, admin access, network segmentation, keyrings, least privilege, encrypted networks, host hardening.

6 lessons
  1. 01A threat model for a Ceph clusterSecurity Hardening · advanced · ~18 min
  2. 02Segmenting and firewalling a Ceph clusterSecurity Hardening · advanced · ~18 min
  3. 03Capability lifecycle: issuance, drift, and reviewSecurity Hardening · advanced · ~18 min
  4. 04Verifying encryption is actually in effectSecurity Hardening · advanced · ~18 min
  5. 05Hardening a Ceph hostSecurity Hardening · advanced · ~18 min
  6. 06Running a repeatable hardening auditSecurity Hardening · advanced · ~18 min

Part CII

Management Security

cephadm SSH, admin keys, dashboard access, monitoring access, REST API authentication, mgr / balancer modules.

6 lessons
  1. 01Securing the cephadm SSH pathManagement Security · advanced · ~18 min
  2. 02Controlling where `client.admin` existsManagement Security · advanced · ~18 min
  3. 03Securing the Ceph dashboardManagement Security · intermediate · ~18 min
  4. 04Securing the monitoring stackManagement Security · intermediate · ~17 min
  5. 05The manager REST API and programmatic accessManagement Security · advanced · ~17 min
  6. 06Manager modules as an attack surfaceManagement Security · advanced · ~17 min

Part CIII

Secrets and Key Management

Keyring protection, client keys, rotation, backup of authentication state, mvcc and metadata db.

6 lessons
  1. 01Finding every copy of a keyringSecrets and Key Management · advanced · ~18 min
  2. 02Distributing keys without copying files aroundSecrets and Key Management · advanced · ~18 min
  3. 03Making key rotation operationally possibleSecrets and Key Management · advanced · ~18 min
  4. 04Backing up and restoring the auth databaseSecrets and Key Management · advanced · ~17 min
  5. 05The monitor store as the cluster's secret storeSecrets and Key Management · advanced · ~18 min
  6. 06Recovering from a key change that broke clientsSecrets and Key Management · advanced · ~18 min

Part CIV

Multi-Tenancy in Practice

RBD pool-isolation, CephFS subvolume-isolation, RGW users and tenants, capability scoping.

6 lessons
  1. 01Tenant onboarding as a repeatable procedureMulti-Tenancy in Practice · advanced · ~18 min
  2. 02Operating CephFS subvolume tenancy over timeMulti-Tenancy in Practice · advanced · ~18 min
  3. 03RGW multi-tenancy in operationMulti-Tenancy in Practice · advanced · ~18 min
  4. 04Namespaces when pool-per-tenant stops scalingMulti-Tenancy in Practice · advanced · ~18 min
  5. 05What happens when a tenant hits their limitMulti-Tenancy in Practice · advanced · ~18 min
  6. 06Continuous isolation verification and cross-tenant incidentsMulti-Tenancy in Practice · advanced · ~18 min

Part CV

Backup Strategy

Ceph redundancy is NOT backup. Replication protects against device and node failure. Backup protects against deletion, corruption, ransomware, operator error.

6 lessons
  1. 01Why replication is not backupBackup Strategy · intermediate · ~17 min
  2. 02Enumerating the threats a backup must answerBackup Strategy · intermediate · ~18 min
  3. 03File-level and object-level backup pathsBackup Strategy · intermediate · ~18 min
  4. 04Snapshots as a control: what they cover and what they costBackup Strategy · intermediate · ~18 min
  5. 05Crash consistency versus application consistencyBackup Strategy · advanced · ~18 min
  6. 06Setting RPO and RTO before designing anythingBackup Strategy · intermediate · ~18 min

Part CVI

RBD Backup

RBD snapshots, rbd export, rbd export-diff, incremental strategies, application consistency.

6 lessons
  1. 01Full RBD export: mechanics and costRBD Backup · intermediate · ~18 min
  2. 02Incremental export with export-diffRBD Backup · advanced · ~18 min
  3. 03Coordinating backup with the guestRBD Backup · advanced · ~18 min
  4. 04Restoring an RBD image from an exportRBD Backup · advanced · ~18 min
  5. 05Rollback versus clone: choosing the recovery shapeRBD Backup · advanced · ~18 min
  6. 06Meeting RBD recovery targets at scaleRBD Backup · advanced · ~18 min

Part CVII

CephFS Backup

CephFS snapshots, file-level backups, subvolume snapshots, recovery validation.

6 lessons
  1. 01CephFS snapshots: the .snap directory and its behaviourCephFS Backup · intermediate · ~18 min
  2. 02Off-cluster file backup with deduplicating toolsCephFS Backup · advanced · ~18 min
  3. 03Subvolume snapshots and asynchronous clonesCephFS Backup · advanced · ~18 min
  4. 04Quiescing shared-filesystem workloadsCephFS Backup · advanced · ~18 min
  5. 05Restoring CephFS dataCephFS Backup · advanced · ~18 min
  6. 06CephFS recovery targets and mirroringCephFS Backup · advanced · ~18 min

Part CVIII

RGW Backup and Replication

Object storage backup, RGW multisite replication, RPO and RTO modelling.

6 lessons
  1. 01Copying buckets with the S3 APIRGW Backup and Replication · intermediate · ~18 min
  2. 02RGW multi-site replication as a DR mechanismRGW Backup and Replication · advanced · ~18 min
  3. 03Building an RGW backup inventoryRGW Backup and Replication · intermediate · ~17 min
  4. 04Bucket configuration that a data copy does not carryRGW Backup and Replication · advanced · ~18 min
  5. 05Versioning and object lock as in-place protectionRGW Backup and Replication · advanced · ~18 min
  6. 06RGW recovery targets and failoverRGW Backup and Replication · advanced · ~18 min

Part CIX

Disaster Recovery

Multiple OSD loss, MON loss, site loss, complete cluster loss. The recovery scenario set.

6 lessons
  1. 01Classifying a failure by what it costsDisaster Recovery · advanced · ~18 min
  2. 02The dependency order of recoveryDisaster Recovery · advanced · ~18 min
  3. 03Scenario modelling: what survives each failureDisaster Recovery · advanced · ~18 min
  4. 04Rebuilding a cluster from nothingDisaster Recovery · advanced · ~18 min
  5. 05Writing a DR plan people can actually executeDisaster Recovery · intermediate · ~18 min
  6. 06Running DR drills that find somethingDisaster Recovery · advanced · ~18 min

Part CX

Monitor Recovery

Rebuilding and recovering MONs safely. monmap, monitor store recovery, hostname changes.

6 lessons
  1. 01The monmap and the cluster identityMonitor Recovery · intermediate · ~17 min
  2. 02Replacing a lost monitorMonitor Recovery · advanced · ~18 min
  3. 03Backing up and restoring the monitor storeMonitor Recovery · advanced · ~18 min
  4. 04What the monitor store holds and how it is actually protectedMonitor Recovery · advanced · ~18 min
  5. 05Moving a monitor to a different addressMonitor Recovery · advanced · ~18 min
  6. 06How a monitor rejoins quorumMonitor Recovery · advanced · ~18 min

Part CXI

Manager Recovery

Active / standby MGR, recover MGR services.

6 lessons
  1. 01The manager active and standby modelManager Recovery · intermediate · ~17 min
  2. 02Replacing a manager daemonManager Recovery · advanced · ~18 min
  3. 03Module state across a manager failoverManager Recovery · advanced · ~18 min
  4. 04Operating a cluster with no managerManager Recovery · advanced · ~18 min
  5. 05Verifying manager stateManager Recovery · intermediate · ~17 min
  6. 06CephFS through a manager outageManager Recovery · advanced · ~18 min

Part CXII

OSD Host Loss

Entire OSD server disappears. Reasoning about CRUSH failure domain and data safety.

6 lessons
  1. 01How the cluster learns a host is goneOSD Host Loss · intermediate · ~18 min
  2. 02Deciding whether to let recovery startOSD Host Loss · intermediate · ~18 min
  3. 03Marking out and the two data movementsOSD Host Loss · advanced · ~18 min
  4. 04How much more you can afford to loseOSD Host Loss · advanced · ~18 min
  5. 05Whether the recovery will fitOSD Host Loss · advanced · ~18 min
  6. 06Controlling how fast recovery runsOSD Host Loss · advanced · ~18 min

Part CXIII

Multiple OSD Failure

Replication size 3 plus multiple OSD losses. Reasoning about remaining replicas.

6 lessons
  1. 01Below min_size and still holding the dataMultiple OSD Failure · advanced · ~18 min
  2. 02When a PG has no surviving copyMultiple OSD Failure · advanced · ~18 min
  3. 03Reading the shape of a correlated failureMultiple OSD Failure · advanced · ~18 min
  4. 04Tolerance is a per-pool propertyMultiple OSD Failure · advanced · ~18 min
  5. 05The imbalance recovery leaves behindMultiple OSD Failure · advanced · ~17 min
  6. 06When more domains fail than the rule allowsMultiple OSD Failure · advanced · ~18 min

Part CXIV

Complete Storage Node Loss

Recovering safely after a host is gone.

6 lessons
  1. 01When the node is never coming backComplete Storage Node Loss · intermediate · ~18 min
  2. 02Removing a host that cannot be drainedComplete Storage Node Loss · advanced · ~18 min
  3. 03How long the rebuild actually takesComplete Storage Node Loss · advanced · ~18 min
  4. 04Whether the rebuild can finish at allComplete Storage Node Loss · advanced · ~18 min
  5. 05Bringing the replacement into the clusterComplete Storage Node Loss · intermediate · ~18 min
  6. 06Two nodes gone at onceComplete Storage Node Loss · advanced · ~18 min

Part CXV

Cluster-Wide Capacity Incident

nearfull, backfillfull, full. Immediate priorities.

6 lessons
  1. 01Finding where the cluster is actually fullCluster-Wide Capacity Incident · intermediate · ~18 min
  2. 02Mapping the blast radiusCluster-Wide Capacity Incident · advanced · ~18 min
  3. 03The first thirty minutesCluster-Wide Capacity Incident · advanced · ~18 min
  4. 04When the imbalance is the emergencyCluster-Wide Capacity Incident · advanced · ~18 min
  5. 05Adding capacity while the cluster is already at the limitCluster-Wide Capacity Incident · advanced · ~18 min
  6. 06Raising the ratios as a bridge, not a fixCluster-Wide Capacity Incident · advanced · ~18 min

Part CXVI

Network Partition

Ceph behaviour during partial network connectivity. MON quorum and OSD communication implications.

6 lessons
  1. 01Reading the first ninety seconds of a partitionNetwork Partition · intermediate · ~18 min
  2. 02One-way reachability and the flap it producesNetwork Partition · advanced · ~18 min
  3. 03Why a partition does not corrupt your dataNetwork Partition · advanced · ~18 min
  4. 04Which network partitioned, and how the symptoms differNetwork Partition · advanced · ~18 min
  5. 05The cost of the partition healingNetwork Partition · advanced · ~18 min
  6. 06Reconstructing the partition from evidenceNetwork Partition · intermediate · ~17 min

Part CXVII

Lost Monitor Quorum

Deep troubleshooting for MON quorum loss, including split-brain recovery.

6 lessons
  1. 01Three monitors and what the arithmetic actually buysLost Monitor Quorum · advanced · ~18 min
  2. 02Five monitors and the placement that makes them countLost Monitor Quorum · advanced · ~18 min
  3. 03Getting one monitor back is the whole jobLost Monitor Quorum · advanced · ~18 min
  4. 04Forcing quorum by editing the monmapLost Monitor Quorum · advanced · ~18 min
  5. 05Triaging why quorum is goneLost Monitor Quorum · advanced · ~18 min
  6. 06The artefacts and the drill that make the recovery routineLost Monitor Quorum · intermediate · ~17 min

Part CXVIII

Data Integrity Incident

Inconsistent PGs, checksum errors, disk errors, repair decisions.

6 lessons
  1. 01Recognising a data integrity incidentData Integrity Incident · advanced · ~18 min
  2. 02Choosing which copy is authoritativeData Integrity Incident · advanced · ~18 min
  3. 03Deciding between repair and restoreData Integrity Incident · advanced · ~18 min
  4. 04When repair does not repairData Integrity Incident · advanced · ~18 min
  5. 05Restoring corrupt data from backupData Integrity Incident · advanced · ~18 min
  6. 06Finding the component that produced the corruptionData Integrity Incident · advanced · ~17 min

Part CXIX

Disaster Recovery Architecture

Primary Ceph cluster to backups and replication to independent recovery domain.

6 lessons
  1. 01Choosing the failure domain a second copy must surviveDisaster Recovery Architecture · advanced · ~18 min
  2. 02Tiering copies for two different recovery timesDisaster Recovery Architecture · intermediate · ~18 min
  3. 03Recovery paths that do not depend on what failedDisaster Recovery Architecture · advanced · ~18 min
  4. 04Measuring the restore time you actually haveDisaster Recovery Architecture · intermediate · ~17 min
  5. 05Encrypting backups and keeping the keys usableDisaster Recovery Architecture · advanced · ~18 min
  6. 06Separating cluster operations from backup operationsDisaster Recovery Architecture · advanced · ~18 min

Part CXX

Multi-Site Concepts

RGW multisite, asynchronous replication concepts, RPO, RTO.

6 lessons
  1. 01What RGW multi-site does that block and file mirroring cannotMulti-Site Concepts · advanced · ~18 min
  2. 02Delivered RPO across three replication mechanismsMulti-Site Concepts · advanced · ~18 min
  3. 03Synchronous cross-site replication and its latency ceilingMulti-Site Concepts · advanced · ~18 min
  4. 04Stretch mode: what enabling it actually changesMulti-Site Concepts · advanced · ~18 min
  5. 05What a replicated copy is consistent withMulti-Site Concepts · advanced · ~18 min
  6. 06Costing a second site, and what to replicate to itMulti-Site Concepts · intermediate · ~17 min

Part CXXI

Storage Architecture Decision-Making

When Ceph is appropriate and when it is not. Workload fit, operational capacity, scale, latency.

6 lessons
  1. 01Testing a workload against Ceph before committing to itStorage Architecture Decision-Making · intermediate · ~17 min
  2. 02Where Ceph is the wrong answerStorage Architecture Decision-Making · intermediate · ~18 min
  3. 03Three hosts: the utilisation ceiling and the repair windowStorage Architecture Decision-Making · intermediate · ~18 min
  4. 04What the sixth and seventh hosts buyStorage Architecture Decision-Making · advanced · ~18 min
  5. 05Ceph on cloud infrastructure: when it is a real choiceStorage Architecture Decision-Making · advanced · ~18 min
  6. 06The architecture decision record and its as-built captureStorage Architecture Decision-Making · intermediate · ~17 min

Part CXXII

Small Cluster Risks

3-node Ceph cluster trade-offs. Recovery bandwidth, maintenance headroom, capacity, fault tolerance.

6 lessons
  1. 01Recovery when there are only two survivorsSmall Cluster Risks · intermediate · ~18 min
  2. 02Maintenance with no spare failure domainSmall Cluster Risks · intermediate · ~17 min
  3. 03Usable capacity when there are few OSDsSmall Cluster Risks · intermediate · ~18 min
  4. 04What three and four hosts actually tolerateSmall Cluster Risks · advanced · ~18 min
  5. 05Patterns that help, and ones that only look like they doSmall Cluster Risks · advanced · ~18 min
  6. 06The triggers that mean another host is dueSmall Cluster Risks · intermediate · ~17 min

Part CXXIII

Capacity and Failure Planning

Normal plus growth plus node failure plus recovery headroom plus maintenance headroom.

6 lessons
  1. 01The capacity budget: what may actually be committedCapacity and Failure Planning · intermediate · ~17 min
  2. 02Growth composition and the real lead timeCapacity and Failure Planning · intermediate · ~18 min
  3. 03Sizing for one host down, or for twoCapacity and Failure Planning · advanced · ~18 min
  4. 04Making a reserve realCapacity and Failure Planning · advanced · ~18 min
  5. 05The maintenance reserve is a scheduling decisionCapacity and Failure Planning · intermediate · ~17 min
  6. 06Working the numbers to an orderCapacity and Failure Planning · advanced · ~18 min

Part CXXIV

Production Reference Architecture

The complete reference design: 6+ nodes, MON/MGR placement, OSD layout, network architecture, services, monitoring, automation, backups.

6 lessons
  1. 01The six-node reference topologyProduction Reference Architecture · advanced · ~18 min
  2. 02Pool and daemon layout for RBD, CephFS, and RGWProduction Reference Architecture · advanced · ~18 min
  3. 03One identity per consumerProduction Reference Architecture · advanced · ~18 min
  4. 04The network the design assumesProduction Reference Architecture · advanced · ~18 min
  5. 05The monitoring stack that ships with the designProduction Reference Architecture · intermediate · ~18 min
  6. 06The design as a repository, and the acceptance testProduction Reference Architecture · advanced · ~18 min

Part Labs

Labs

Hands-on disposable-cluster labs (cephadm build, OSD lifecycle, RBD, CephFS, RGW, recovery, capacity, Proxmox, Kubernetes CSI, network impairment).

0 lessons

No lessons published in this part yet. The full curriculum is planned in docs/courses/ceph/curriculum.md on GitHub.

Part Runbooks

Runbooks

Operational procedures for build, OSD replacement, MON recovery, recovery tuning, capacity incidents, DR, network maintenance.

0 lessons

Part Checklists

Checklists

Production readiness, hardware readiness, network readiness, CRUSH review, pool design, capacity planning, security, RBD/CephFS/RGW readiness, observability, OSD replacement, node maintenance, pre-upgrade, post-upgrade, backup DR, Proxmox + Ceph readiness, Kubernetes + Ceph readiness.

0 lessons

Part Breakfix

Break/Fix Scenarios

Evidence-first diagnosis of Ceph failure modes (OSD, MON, MDS, RBD, CephFS, RGW, network, capacity, recovery, upgrades).

0 lessons

Part Capstone

Capstone: Mission-Critical Ceph

A complete production Ceph estate end to end: 4-node cluster with RBD/CephFS/RGW, Proxmox and Kubernetes consumers, full observability, backup, and 12 controlled-failure incidents.

0 lessons

No lessons published in this part yet. The full curriculum is planned in docs/courses/ceph/curriculum.md on GitHub.

Part Final

Final Assessment

Final theory assessment plus final practical assessment of an inherited Ceph estate.

0 lessons

No lessons published in this part yet. The full curriculum is planned in docs/courses/ceph/curriculum.md on GitHub.