← All runbooks in Observability
Runbook: Observability Platform Outage During Incident
1 · Prerequisites
Confirm every item is in place before any state change.
- Observability platform
- Incident response plan
2 · Pre-checks
Read-only diagnostic commands. If any of these don't match expected output, stop and investigate further.
- · Treat as incident
- · Engage platform team
3 · Procedure
Execute each step in order. Verify the expected output of a step before moving to the next.
- 1Use alternative evidence: saved dashboards, application logs, knowledge
- 2Recover observability platform
- 3Restore configuration from version control
- 4Recreate receivers
- 5Validate
4 · Verification
Confirm the procedure actually fixed the problem.
- ✓The platform is back online
- ✓Recovery is documented
5 · Rollback
If verification fails, undo the procedure in reverse order.
- ↶N/A
6 · Escalation
When the runbook isn't enough, contact:
- · Engage vendor support
- · Practice the DR procedure quarterly
Purpose
Observability Platform Outage During Incident
When to use this runbook
Use this runbook when the operator needs a guided procedure to handle the situation described above.
Pre-checks
Before starting the procedure, confirm the prerequisites and pre-checks are met. The structured lists are rendered from the frontmatter by the page layout.
Procedure
Follow the steps from the frontmatter procedure steps. The page layout renders the steps as a checklist with copy-to-clipboard affordances.
Verification
After the procedure, the structured verification items from the frontmatter are rendered as a checklist.
Rollback
If the procedure fails or makes things worse, follow the structured rollback steps from the frontmatter.
Escalation
The structured escalation path is rendered from the frontmatter. Use it if the operator cannot complete the procedure safely.