Skip to main content
RunBook Academy

← All runbooks in Observability

critical riskcluster affecting~30 min

Runbook: Restore Observability Configuration

1 · Prerequisites

Confirm every item is in place before any state change.

  • Observability backup
  • Configuration in version control

2 · Pre-checks

Read-only diagnostic commands. If any of these don't match expected output, stop and investigate further.

  • · Determine the failure scope
  • · Locate the backup

3 · Procedure

Execute each step in order. Verify the expected output of a step before moving to the next.

  1. 1Restore configuration files from backup
  2. 2Restart affected services
  3. 3Verify /-/ready and dashboards

4 · Verification

Confirm the procedure actually fixed the problem.

  • Configuration is restored
  • Services are running
  • Validation passes

5 · Rollback

If verification fails, undo the procedure in reverse order.

  • Re-run if failed and document

6 · Escalation

When the runbook isn't enough, contact:

  • · Escalate immediately

Purpose

Restore Observability Configuration

When to use this runbook

Use this runbook when the operator needs a guided procedure to handle the situation described above.

Pre-checks

Before starting the procedure, confirm the prerequisites and pre-checks are met. The structured lists are rendered from the frontmatter by the page layout.

Procedure

Follow the steps from the frontmatter procedure steps. The page layout renders the steps as a checklist with copy-to-clipboard affordances.

Verification

After the procedure, the structured verification items from the frontmatter are rendered as a checklist.

Rollback

If the procedure fails or makes things worse, follow the structured rollback steps from the frontmatter.

Escalation

The structured escalation path is rendered from the frontmatter. Use it if the operator cannot complete the procedure safely.

References

  1. Prometheus documentation
  2. Grafana documentation
  3. Loki documentation
  4. Tempo documentation