Architecture
The student examines the broken architecture diagram and identifies the responsible component. The cluster is running Ceph Tentacle (20.2.x); the incident is one a production operator must diagnose from the evidence presented.
Symptoms
- MDS comes back up but takes minutes to serve
- Clients see slow file I/O during replay
- CephFS latency is elevated for an extended period
Evidence
- ceph mds metadata shows a large journal file
- ceph mds status reports replay for an extended time
- ceph daemonperf mds.
<id>shows replay activity
Student investigation
The student follows the methodology: define the symptom, determine the impact, gather evidence, identify the component, form a hypothesis, test safely, restore, validate.
Progressive hints
- Hint 1: check the cluster state with
ceph -sfirst. - Hint 2: read
ceph health detailand identify the affected PGs / OSDs / daemons. - Hint 3: use
ceph osd tree,ceph pg dump, orceph mds statas the next-level diagnostic. - Hint 4: the root cause is documented in the frontmatter
root_causefield.
Validation
The student runs the verification steps and confirms the symptom cleared.
Root cause
See the frontmatter root_cause field.
Remediation
Apply the fix from the frontmatter remediation field.
Prevention
Apply the prevention measures from the frontmatter prevention field.