Architecture
The student examines the broken architecture diagram and identifies the responsible component. The cluster is running Ceph Tentacle (20.2.x); the incident is one a production operator must diagnose from the evidence presented.
Symptoms
- ceph -s shows HEALTH_WARN with osds down
- PGs in active+clean+remapped state
- One OSD host reports I/O errors
- No alert email yet (within SLA window)
Evidence
- ceph osd tree shows osd.5 down with host ceph01
- ceph osd df tree shows osd.5 weight 0 but capacity still reported
- journalctl -u ceph-osd@5 shows OSD crashed 4 hours ago
- dmesg shows I/O errors against /dev/sdc
Student investigation
The student follows the methodology: define the symptom, determine the impact, gather evidence, identify the component, form a hypothesis, test safely, restore, validate.
Progressive hints
- Hint 1: check the cluster state with
ceph -sfirst. - Hint 2: read
ceph health detailand identify the affected PGs / OSDs / daemons. - Hint 3: use
ceph osd tree,ceph pg dump, orceph mds statas the next-level diagnostic. - Hint 4: the root cause is documented in the frontmatter
root_causefield.
Validation
The student runs the verification steps and confirms the symptom cleared.
Root cause
See the frontmatter root_cause field.
Remediation
Apply the fix from the frontmatter remediation field.
Prevention
Apply the prevention measures from the frontmatter prevention field.