Architecture
The student examines the broken architecture diagram and identifies the responsible component. The cluster is running Ceph Tentacle (20.2.x); the incident is one a production operator must diagnose from the evidence presented.
Symptoms
- ceph -s reports multiple OSDs down + in
- Several PGs are active+degraded+undersized
- Two or more OSDs per affected PG are unavailable
Evidence
- ceph osd tree shows which OSDs are down
- ceph pg dump_stuck shows the affected PGs
- ceph health detail lists recovery_too_slow and reduced_data_redundancy
Student investigation
The student follows the methodology: define the symptom, determine the impact, gather evidence, identify the component, form a hypothesis, test safely, restore, validate.
Progressive hints
- Hint 1: check the cluster state with
ceph -sfirst. - Hint 2: read
ceph health detailand identify the affected PGs / OSDs / daemons. - Hint 3: use
ceph osd tree,ceph pg dump, orceph mds statas the next-level diagnostic. - Hint 4: the root cause is documented in the frontmatter
root_causefield.
Validation
The student runs the verification steps and confirms the symptom cleared.
Root cause
See the frontmatter root_cause field.
Remediation
Apply the fix from the frontmatter remediation field.
Prevention
Apply the prevention measures from the frontmatter prevention field.