Observability · Operational reference
Toolkit
Operational artefacts for Observability: checklists for cadences, runbooks for incidents, hands-on labs for skills, and break/fix scenarios for troubleshooting reflexes. None of this replaces reading the corresponding lessons — these are the artefacts you keep open in a second tab during real work.
- Runbooks
- 30
- Checklists
- 16
- Labs
- 30
- Break/Fix
- 32
- Assessments
- 1
Operational checklists
Repetitive tasks performed on a schedule.
Before deployment
- Alerting Readiness21 items→
- Backup and DR Readiness22 items→
- Grafana Production Readiness23 items→
- Loki Production Readiness24 items→
- OpenTelemetry Collector Readiness29 items→
- Observability Platform Production Readiness12 items→
- Prometheus Production Readiness29 items→
- Telemetry Security Review24 items→
- Tempo Production Readiness25 items→
Quarterly
Runbooks
Step-by-step procedures. Treat the rollback as part of the procedure — never skip it.
Critical risk
- criticaldata loss riskRunbook: Recover Loki~90 min · 14 steps · verified not-executed→
- criticaldata loss riskRunbook: Full Observability DR~240 min · 15 steps · verified 2026-08-18→
- criticalcluster affectingRunbook: Observability Platform Outage During Incident~45 min · 13 steps · verified 2026-08-19→
- criticalcluster affectingRunbook: Restore Observability Configuration~60 min · 14 steps · verified 2026-08-19→
- criticalcluster affectingRunbook: Recover a Broken Prometheus~90 min · 12 steps · verified 2026-08-19→
- criticaldata loss riskRunbook: Respond to a Telemetry Data Leak~240 min · 13 steps · verified 2026-08-19→
- criticaldata loss riskRunbook: Recover Tempo~90 min · 16 steps · verified not-executed→
High risk
- highservice affectingRunbook: Investigate an Alert Not Firing~30 min · 11 steps · verified 2026-08-18→
- highservice affectingRunbook: Investigate Alertmanager Delivery Failure~30 min · 12 steps · verified 2026-08-18→
- highcluster affectingRunbook: Investigate High Cardinality~45 min · 12 steps · verified 2026-08-18→
- highcluster affectingRunbook: Investigate Loki High Ingestion Volume~45 min · 12 steps · verified 2026-08-18→
- highcluster affectingRunbook: Investigate Loki Ingestion Failure~45 min · 12 steps · verified 2026-08-18→
- highcluster affectingRunbook: Observability Storage Full~60 min · 13 steps · verified 2026-08-19→
- highservice affectingRunbook: Upgrade the Full Stack~240 min · 12 steps · verified 2026-08-19→
- highservice affectingRunbook: Investigate OTel Collector Failure~45 min · 12 steps · verified 2026-08-19→
- highcluster affectingRunbook: Investigate Prometheus High Memory~45 min · 12 steps · verified 2026-08-19→
- highcluster affectingRunbook: Investigate Tempo Ingestion Failure~60 min · 13 steps · verified 2026-08-19→
- highservice affectingRunbook: Certificate Expiry Incident~60 min · 13 steps · verified not-executed→
Medium risk
- mediumservice affectingRunbook: Investigate a Correlation Failure~30 min · 11 steps · verified not-executed→
- mediumservice affectingRunbook: Investigate a Failed Grafana Data Source~30 min · 10 steps · verified not-executed→
- mediumservice affectingRunbook: Install Grafana~30 min · 11 steps · verified not-executed→
- mediumservice affectingRunbook: Investigate a Loki Query Failure~40 min · 12 steps · verified not-executed→
- mediumservice affectingRunbook: Investigate a Missing Metric~35 min · 13 steps · verified not-executed→
- mediumservice affectingRunbook: Investigate a Missing Prometheus Target~30 min · 14 steps · verified 2026-08-18→
- mediuminformationalRunbook: Install Prometheus~90 min · 11 steps · verified 2026-08-19→
- mediumservice affectingRunbook: Upgrade Prometheus~120 min · 12 steps · verified 2026-08-19→
- mediumservice affectingRunbook: Restore Grafana Dashboards and Configuration~90 min · 13 steps · verified 2026-08-19→
- mediumservice affectingRunbook: Investigate Missing Traces~45 min · 13 steps · verified not-executed→
Hands-on labs
Time-boxed exercises.
B · Nested virtualisation
- Lab: Build an Alert Rule~90 min · 5 objectives→
- Lab: Configure Alertmanager~90 min · 5 objectives→
- Lab: Backup and Restore~90 min · 5 objectives→
- Lab: Correlate Metrics, Logs, and Traces~90 min · 5 objectives→
- Lab: Deploy Grafana~75 min · 5 objectives→
- Lab: Provision Grafana from Disk~75 min · 5 objectives→
- Lab: Design HA~90 min · 5 objectives→
- Lab: Histograms and Latency~75 min · 5 objectives→
- Lab: Identify Loki Cardinality Incidents~90 min · 5 objectives→
- Lab: Query Logs with LogQL~90 min · 5 objectives→
- Lab: Deploy Loki~90 min · 5 objectives→
- Lab: Investigate a Metric -> Logs -> Trace Workflow~90 min · 5 objectives→
- Lab: Derive Metrics from Logs~90 min · 5 objectives→
- Lab: Run node_exporter on Linux~90 min · 5 objectives→
- Lab: Build an OpenTelemetry Collector Pipeline~90 min · 5 objectives→
- Lab: Capstone Production Observability Estate~180 min · 5 objectives→
- Lab: Build a Production Dashboard~90 min · 5 objectives→
- Lab: Deploy Prometheus~90 min · 5 objectives→
- Lab: PromQL Selectors and Matchers~70 min · 5 objectives→
- Lab: Rate and Aggregation~80 min · 5 objectives→
- Lab: Build a Recording Rule~85 min · 5 objectives→
- Lab: SLO-Based Alerting~90 min · 5 objectives→
- Lab: Ship Structured Logs~90 min · 5 objectives→
- Lab: Secure the Observability Stack~90 min · 5 objectives→
- Lab: Deploy Tempo~90 min · 5 objectives→
- Lab: Query Traces with TraceQL~75 min · 5 objectives→
- Lab: Instrument and Sample Traces~90 min · 5 objectives→
Break/Fix scenarios
Troubleshooting drills. Each scenario gives symptoms and evidence, then hides the solution behind a reveal.
prometheus-rules
alertmanager
metric-cardinality
certificate-expiry
grafana-dashboard
loki-ingestion
observability-of-observability
prometheus-tsdb
prometheus-scrape
observability-security
tempo-propagation
Assessments
Production-readiness examinations. Each combines auto-scored questions with scenario-based rubrics you can use to grade yourself.