Observability · Operational reference
Toolkit
Operational artefacts for Observability: checklists for cadences, runbooks for incidents, hands-on labs for skills, and break/fix scenarios for troubleshooting reflexes. None of this replaces reading the corresponding lessons — these are the artefacts you keep open in a second tab during real work.
- Runbooks
- 30
- Checklists
- 16
- Labs
- 30
- Break/Fix
- 32
- Assessments
- 1
Operational checklists
Repetitive tasks performed on a schedule.
Before deployment
- Alerting Readiness12 items→
- Backup and DR Readiness12 items→
- Grafana Production Readiness12 items→
- Loki Production Readiness12 items→
- OpenTelemetry Collector Readiness12 items→
- Observability Platform Production Readiness12 items→
- Prometheus Production Readiness12 items→
- Telemetry Security Review12 items→
- Tempo Production Readiness12 items→
Quarterly
Runbooks
Step-by-step procedures. Treat the rollback as part of the procedure — never skip it.
Critical risk
- criticalcluster affectingRunbook: Recover Loki~30 min · 5 steps · verified 2026-08-13→
- criticaldata loss riskRunbook: Full Observability DR~30 min · 6 steps · verified 2026-08-13→
- criticalcluster affectingRunbook: Observability Platform Outage During Incident~30 min · 5 steps · verified 2026-08-13→
- criticalcluster affectingRunbook: Restore Observability Configuration~30 min · 3 steps · verified 2026-08-13→
- criticalcluster affectingRunbook: Recover a Broken Prometheus~30 min · 7 steps · verified 2026-08-13→
- criticaldata loss riskRunbook: Respond to a Telemetry Data Leak~30 min · 6 steps · verified 2026-08-13→
- criticalcluster affectingRunbook: Recover Tempo~30 min · 5 steps · verified 2026-08-13→
High risk
- highservice affectingRunbook: Investigate an Alert Not Firing~30 min · 7 steps · verified 2026-08-13→
- highservice affectingRunbook: Investigate Alertmanager Delivery Failure~30 min · 7 steps · verified 2026-08-13→
- highservice affectingRunbook: Investigate a Correlation Failure~30 min · 5 steps · verified 2026-08-13→
- highcluster affectingRunbook: Investigate High Cardinality~30 min · 6 steps · verified 2026-08-13→
- highcluster affectingRunbook: Investigate Loki High Ingestion Volume~30 min · 6 steps · verified 2026-08-13→
- highcluster affectingRunbook: Investigate Loki Ingestion Failure~30 min · 7 steps · verified 2026-08-13→
- highcluster affectingRunbook: Observability Storage Full~30 min · 7 steps · verified 2026-08-13→
- highservice affectingRunbook: Upgrade the Full Stack~30 min · 8 steps · verified 2026-08-13→
- highservice affectingRunbook: Investigate OTel Collector Failure~30 min · 5 steps · verified 2026-08-13→
- highcluster affectingRunbook: Investigate Prometheus High Memory~30 min · 6 steps · verified 2026-08-13→
- highcluster affectingRunbook: Investigate Tempo Ingestion Failure~30 min · 6 steps · verified 2026-08-13→
- highservice affectingRunbook: Investigate Missing Traces~30 min · 6 steps · verified 2026-08-13→
- highservice affectingRunbook: Certificate Expiry Incident~30 min · 6 steps · verified 2026-08-13→
Medium risk
- mediumservice affectingRunbook: Investigate a Failed Grafana Data Source~30 min · 6 steps · verified 2026-08-13→
- mediumservice affectingRunbook: Install Grafana~30 min · 6 steps · verified 2026-08-13→
- mediuminformationalRunbook: Investigate a Loki Query Failure~30 min · 6 steps · verified 2026-08-13→
- mediumservice affectingRunbook: Investigate a Missing Metric~30 min · 5 steps · verified 2026-08-13→
- mediumservice affectingRunbook: Investigate a Missing Prometheus Target~30 min · 6 steps · verified 2026-08-13→
- mediumservice affectingRunbook: Install Prometheus~30 min · 7 steps · verified 2026-08-13→
- mediumservice affectingRunbook: Upgrade Prometheus~30 min · 6 steps · verified 2026-08-13→
- mediumservice affectingRunbook: Restore Grafana Dashboards and Configuration~30 min · 5 steps · verified 2026-08-13→
Hands-on labs
Time-boxed exercises.
C · Simulation
- Lab: Build an Alert Rule~60 min · 4 objectives→
- Lab: Configure Alertmanager~60 min · 4 objectives→
- Lab: Backup and Restore~60 min · 4 objectives→
- Lab: Blackbox Probes~60 min · 4 objectives→
- Lab: Cardinality Optimisation~60 min · 4 objectives→
- Lab: Recover from a Cardinality Incident~60 min · 4 objectives→
- Lab: Correlate Metrics, Logs, and Traces~60 min · 4 objectives→
- Lab: Deploy Grafana~60 min · 4 objectives→
- Lab: Provision Grafana from Disk~60 min · 4 objectives→
- Lab: Design HA~60 min · 4 objectives→
- Lab: Histograms and Latency~60 min · 4 objectives→
- Lab: Identify Loki Cardinality Incidents~60 min · 4 objectives→
- Lab: Query Logs with LogQL~60 min · 4 objectives→
- Lab: Deploy Loki~60 min · 4 objectives→
- Lab: Investigate a Metric -> Logs -> Trace Workflow~60 min · 4 objectives→
- Lab: Derive Metrics from Logs~60 min · 4 objectives→
- Lab: Run node_exporter on Linux~60 min · 4 objectives→
- Lab: Build an OpenTelemetry Collector Pipeline~60 min · 4 objectives→
- Lab: Capstone Production Observability Estate~60 min · 4 objectives→
- Lab: Build a Production Dashboard~60 min · 4 objectives→
- Lab: Deploy Prometheus~60 min · 4 objectives→
- Lab: PromQL Selectors and Matchers~60 min · 4 objectives→
- Lab: Rate and Aggregation~60 min · 4 objectives→
- Lab: Build a Recording Rule~60 min · 4 objectives→
- Lab: SLO-Based Alerting~60 min · 4 objectives→
- Lab: Ship Structured Logs~60 min · 4 objectives→
- Lab: Secure the Observability Stack~60 min · 4 objectives→
- Lab: Deploy Tempo~60 min · 4 objectives→
- Lab: Query Traces with TraceQL~60 min · 4 objectives→
- Lab: Instrument and Sample Traces~60 min · 4 objectives→
Break/Fix scenarios
Troubleshooting drills. Each scenario gives symptoms and evidence, then hides the solution behind a reveal.
alertmanager
metric-cardinality
certificate-expiry
prometheus-rules
grafana-dashboard
loki-ingestion
observability-of-observability
prometheus-tsdb
prometheus-scrape
observability-security
tempo-propagation
observability-scaling
Assessments
Production-readiness examinations. Each combines auto-scored questions with scenario-based rubrics you can use to grade yourself.