Skip to main content
RunBook Academy

observability · monitoring

Observability for Production Sysadmins

A hands-on, fully visual course that takes a systems administrator from "I have a few metrics in a dashboard" to "I can take operational responsibility for a production metrics / logs / traces platform and use it to answer the questions that matter during incidents." Covers the principles of observability vs monitoring, Prometheus architecture, exporters, PromQL, alerting and Alertmanager, SLO-based alerting, Grafana dashboards and provisioning, Loki labels and cardinality, LogQL, Tempo tracing and TraceQL, OpenTelemetry, metric / log / trace correlation, security, scalability, capacity, retention, configuration as code, CI validation, monitoring-the-monitoring-stack, backup, DR, upgrades, incident investigation workflows, and a capstone production observability estate.

Who this is for

  • Linux systems administrators owning the observability platform
  • Infrastructure / platform engineers running Prometheus / Grafana / Loki / Tempo
  • SREs and DevOps engineers responsible for production telemetry
  • Security engineers reviewing observability deployments
  • Cloud engineers provisioning distributed observability infrastructure

Prerequisites

  • Comfortable administering Linux from the shell (systemd, journalctl, filesystems, networking)
  • Working knowledge of Docker and Docker Compose (the labs run as Compose stacks)
  • Basic familiarity with HTTP, TLS, DNS, and TCP/IP
  • Some experience running Prometheus or Grafana is helpful but not required

Other RunBook Academy courses

  • Linux — recommended. Linux is where most telemetry comes from: node_exporter, journald, syslog, NTP, SELinux. The RunBook Academy Linux course covers every host-side primitive this course then exposes to Prometheus and Loki.
  • Docker & Containers — recommended. The Observability labs run as Docker Compose stacks: Prometheus, Grafana, Loki, Tempo, Grafana Alloy, exporters, and OpenTelemetry Collector all run as containers. The Docker course teaches the lifecycle they all share.
  • Ansible — recommended. Ansible is the standard configuration-management layer for a production observability platform: prometheus.yml, alertmanager.yml, Grafana provisioning, Loki and Tempo configurations. The Ansible course teaches the patterns this course then applies.
  • Terraform — recommended. Terraform is the optional provisioning layer for observability control-plane resources (remote state, lock tables, object storage, IAM). The Terraform course teaches the state-driven IaC patterns referenced in the production architecture parts.

What you'll be able to do

After completing this course, you should be capable of independently:

  • Explain the difference between monitoring and observability, and the role of metrics, logs, and traces in operating production systems
  • Design SLIs, SLOs, and error budgets appropriate to a business service
  • Identify and avoid dangerous label cardinality in metrics and logs
  • Deploy, configure, secure, and operate Prometheus in production
  • Write production PromQL: rates, aggregation, histograms, vector matching, recording rules
  • Design useful alerts that fire on user impact and avoid alert fatigue
  • Configure Alertmanager routing, grouping, inhibition, silences, and receivers
  • Deploy, configure, and operate Grafana including provisioning of dashboards as code
  • Operate Loki safely: stream labels, retention, storage, query performance
  • Write production LogQL: stream selectors, parsers, label filters, log-derived metrics
  • Deploy Tempo, ship traces, query with TraceQL, and reason about sampling
  • Use OpenTelemetry Collector as a vendor-neutral telemetry pipeline
  • Correlate metrics, logs, and traces to investigate incidents
  • Monitor a Linux fleet, Docker workloads, network services, databases, and applications
  • Design and validate production architecture: HA, storage, retention, capacity, cost
  • Secure telemetry: TLS, auth, RBAC, sensitive data redaction, multi-tenancy
  • Manage observability configuration as code, with CI validation and rule testing
  • Monitor the monitoring platform, design meta-monitoring, and escape recursive dependencies
  • Back up, recover, and upgrade the observability platform safely
  • Investigate production incidents using evidence from all three signals
  • Complete a capstone: a production observability estate for a Linux + Docker + PostgreSQL workload, with failure injection and recovery

Curriculum overview

120 planned parts · 684 lessons currently published.

Part I

Foundations

Monitoring vs observability, telemetry signals, symptoms vs causes, evidence-driven investigation.

6 lessons

Part II

Production Monitoring Fundamentals

SLIs, SLOs, error budgets, USE and RED methodologies, where the limits of each lie.

6 lessons

Part III

Metrics Fundamentals

Counters, gauges, histograms, summaries, labels, dimensions, time series, aggregation.

6 lessons

Part IV

Cardinality

Label design, series estimation, dangerous labels, cardinality budgets, cardinality incidents.

6 lessons

Part V

Prometheus Architecture

Pull model, scrape lifecycle, TSDB, rules, Alertmanager, exporters, federation, remote write.

6 lessons

Part VI

Installing Prometheus

Linux binaries, systemd, containers, Docker Compose for labs, storage path, retention, permissions.

6 lessons

Part VII

Prometheus Configuration

global, scrape_configs, rule_files, alerting, remote_write, interval semantics.

6 lessons

Part VIII

Service Discovery

Static, file-based, cloud, container, DNS — with relabelling and target management.

6 lessons

Part IX

Exporters

Exporter contract, node_exporter, blackbox_exporter, application and database exporters.

6 lessons

Part X

node_exporter

Collectors and metrics for CPU, memory, filesystems, disks, network, load.

6 lessons

Part XI

Blackbox Monitoring

HTTP, TCP, ICMP, DNS, TLS — synthetic probing and what white-box cannot see.

6 lessons

Part XII

PromQL Foundations

Selectors, label matchers, range and instant vectors, scalars, functions, operators.

6 lessons

Part XIII

Rates and Counters

rate, irate, increase, counter resets, why raw counter graphs mislead.

6 lessons

Part XIV

Aggregation

sum, avg, min, max, count, topk, bottomk, by, without — operational interpretation.

6 lessons

Part XV

Histograms and Latency

Buckets, _bucket/_sum/_count, quantiles, native histograms, percentile maths.

6 lessons

Part XVI

PromQL Troubleshooting

Wrong rate window, missing labels, vector matching, stale series, over-aggregation.

6 lessons

Part XVII

Recording Rules

Expensive PromQL into reuse; naming conventions, evaluation, performance, maintenance.

6 lessons

Part XVIII

Alerting Rules

Metric → condition → for → alert; labels, annotations, severity, ownership.

6 lessons

Part XIX

Alertmanager

Routing, grouping, inhibition, silences, receivers, deduplication, lifecycle.

6 lessons

Part XX

Alert Quality

Symptoms vs causes, page vs ticket, alert fatigue, severity, ownership.

6 lessons

Part XXI

Alert Inhibition

Dependency-aware alerting, host-down suppresses exporter and downstream alerts.

6 lessons

Part XXII

SLO-Based Alerting

Error-budget burn-rate, multi-window alerting, when SRE maths is the right tool.

6 lessons

Part XXIII

Grafana Foundations

Organisations, users, data sources, dashboards, panels, variables, provisioning model.

6 lessons

Part XXIV

Grafana Installation

Package, container, storage, configuration, systemd, reverse proxy, TLS, auth.

6 lessons

Part XXV

Grafana Data Sources

Prometheus, Loki, Tempo — and the resource-id correlation they enable.

6 lessons

Part XXVI

Dashboard Design

Overview vs detail, units, thresholds, variables, links, annotations, ownership.

6 lessons

Part XXVII

Dashboard Anti-Patterns

Wallpaper dashboards, rainbow, meaningless gauges, inconsistent units, no context, no link.

6 lessons

Part XXVIII

Grafana Variables

Template variables, repeated panels, environment, region, instance, query cost.

6 lessons

Part XXIX

Grafana Provisioning

Dashboards and data sources as code, YAML schema, Git workflow, drift avoidance.

6 lessons

Part XXX

Grafana Security

Auth, RBAC, anonymous access, secrets, plugins, reverse proxy, TLS, session.

6 lessons

Part XXXI

Logging Foundations

Structured vs unstructured, severity, timestamps, correlation IDs, machine readability.

6 lessons

Part XXXII

Logging Pipeline Architecture

Agent / collector / ingest / index / query; the strategic choice between Grafana Alloy and OTel Collector.

6 lessons

Part XXXIII

Loki Architecture

Streams, labels, chunks, indexes, ingesters, queriers, compactor, retention, storage.

6 lessons

Part XXXIV

Loki Labels and Cardinality

Low-cardinality dimensions only; high-cardinality values stay in log content.

6 lessons

Part XXXV

Loki Installation

Small installs to production architecture; storage, retention, limits, ingestion.

6 lessons

Part XXXVI

Log Shipping

Grafana Alloy as the strategic log collector; OpenTelemetry Collector as the vendor-neutral alternative; Promtail only for legacy environments.

6 lessons

Part XXXVII

LogQL Foundations

Stream selectors, line filters, parsers, label filters, regex, structured logs.

6 lessons

Part XXXVIII

LogQL Metrics

Derive metrics from logs; error rates, log volume, log-derived latency.

6 lessons

Part XXXIX

Log Troubleshooting

Missing logs, timestamps, label mismatch, ingestion, parser errors, query performance.

6 lessons

Part XL

Log Retention

Retention configuration, capacity, legal/security considerations, deletion, cost.

6 lessons

Part XLI

Distributed Tracing Foundations

Trace, span, parent / child, trace ID, span ID, attributes, events, status.

6 lessons

Part XLII

Why Tracing Exists

What traces answer that metrics and logs often cannot; dependency latency, execution path.

6 lessons

Part XLIII

Instrumentation

Manual / automatic instrumentation, OTel SDK, context propagation, attributes.

6 lessons

Part XLIV

Sampling

Head sampling, tail sampling, sampling rate, rare errors, latency, cost.

6 lessons

Part XLV

Tempo Architecture

Distributor, ingester, querier, compactor, metrics-generator, object storage.

6 lessons

Part XLVI

Tempo Deployment

Single-binary vs microservices, storage backend, receivers, retention.

6 lessons

Part XLVII

Trace Queries

TraceQL — selectors, filters, aggregations, spans, attributes, intrisics.

6 lessons

Part XLVIII

Trace Troubleshooting

Missing spans, propagation, sampling, clock, collector, ingest, attributes.

6 lessons

Part XLIX

OpenTelemetry Foundations

Vendor-neutral telemetry: signals, SDK, OTel Collector, OTLP, Agent and Gateway patterns.

6 lessons

Part L

OpenTelemetry Collector

Receivers, processors, exporters, pipelines, extensions, resource detection.

6 lessons

Part LI

Correlating Metrics, Logs, and Traces

Alert → dashboard → metrics → logs → traces → root cause, end-to-end.

6 lessons

Part LII

Exemplars

Linking metric datapoints to traces; the operational value of an exemplar trail.

6 lessons

Part LIII

Log / Trace Correlation

Trace IDs in structured logs, service metadata, jump-to-trace from logs.

6 lessons

Part LIV

Dashboard-to-Logs Workflows

Symptom in a panel → relevant log stream; correlation drilling.

6 lessons

Part LV

Dashboard-to-Traces Workflows

Latency / error metric → exemplar → trace; finding the slow dependency.

6 lessons

Part LVI

Linux Observability

Host metrics, filesystem metrics, network metrics, journal, syslog, time sync, cluster-level.

6 lessons

Part LVII

Docker Observability

Host + container + application metrics; cAdvisor; short-lived workloads; file logs.

6 lessons

Part LVIII

Proxmox Observability

Cluster health, VM / LXC resource utilisation, storage, network, backup signals.

6 lessons

Part LIX

Database Observability

Connections, query latency, transactions, locks, cache, replication, storage.

6 lessons

Part LX

Network Observability

Interface, errors, drops, latency, availability, blackbox probes, DNS, ICMP.

6 lessons

Part LXI

Application Observability

RED, request rate, error rate, duration, saturation, custom metrics, dependency performance.

6 lessons

Part LXII

Business Metrics

When business-level signals improve operations; observability vs analytics.

6 lessons

Part LXIII

Synthetic Monitoring

External probes; DNS, TCP, TLS, HTTP, multi-step journeys, user-perspective checks.

6 lessons

Part LXIV

TLS Monitoring

Certificate expiry, handshake success, protocol health, expiry incident scenarios.

6 lessons

Part LXV

DNS Monitoring

Resolution, slow resolution, wrong record, stale record, external vs internal.

6 lessons

Part LXVI

Observability Architecture for Production

Production → collectors → metrics/logs/traces → Prometheus/Loki/Tempo → Grafana → Alertmanager.

6 lessons

Part LXVII

High Availability

Stateless vs stateful components; HA is a property of the system, not the deployment.

6 lessons

Part LXVIII

Prometheus HA

Duplicate scraping, external labels, replica quorum, Active-Active, Active-Passive.

6 lessons

Part LXIX

Long-Term Metrics Storage

Why single-node Prometheus has limits; remote write; Thanos / Mimir concepts; when they become necessary.

6 lessons

Part LXX

Loki at Scale

Ingest, query, storage, caches, object storage, replication, limits, schema config.

6 lessons

Part LXXI

Tempo at Scale

Ingestion, object storage, scaling, retention, compaction, query performance.

6 lessons

Part LXXII

Grafana HA

Shared database, sessions, provisioning, load balancing, plugins, configuration consistency.

6 lessons

Part LXXIII

Storage Architecture

Local disks vs object storage; throughput, IOPS, retention, capacity for the observability platform.

6 lessons

Part LXXIV

Capacity Planning

Samples / sec, active series, log ingestion rate, trace ingestion rate, retention, growth.

6 lessons

Part LXXV

Performance

Identifying bottlenecks inside the observability platform itself; scrape overload, query pressure.

6 lessons

Part LXXVI

Cost Management

Series × samples × retention; bytes / sec × retention; spans / sec × sampling.

6 lessons

Part LXXVII

Security Architecture

Network, auth, TLS, secrets, multi-tenancy, sensitive data, tenant isolation.

6 lessons

Part LXXVIII

Securing Prometheus

Network exposure, basic auth, admin endpoints, exporters, TLS where appropriate.

6 lessons

Part LXXIX

Securing Grafana

Admin, SSO, teams, RBAC, data-source credentials, anonymous, plugins, secrets.

6 lessons

Part LXXX

Securing Loki

Log confidentiality, tenant isolation, API access, object-store permissions, log injection.

6 lessons

Part LXXXI

Securing Tempo

Trace data sensitivity, access control, object storage, attributes containing PII or secrets.

6 lessons

Part LXXXII

Secrets and Sensitive Telemetry

Observability often captures what production should never log; passwords, tokens, PII, payment.

6 lessons

Part LXXXIII

Multi-Tenancy

Organisational / tenant separation; teams, environments, customers, access boundaries.

6 lessons

Part LXXXIV

Configuration as Code

Prometheus, alerting, Alertmanager, Grafana provisioning, Loki, Tempo, collectors in Git.

6 lessons

Part LXXXV

CI Validation

Config change → syntax check → rule check → lint → test → review → deploy.

6 lessons

Part LXXXVI

Prometheus Rule Testing

promtool check rules, unit tests with synthetic series; rule tests as regression guards.

6 lessons

Part LXXXVII

Alert Testing

Testing alerts before production; synthetic series, receiver tests, canary alerts.

6 lessons

Part LXXXVIII

Dashboard Testing and Review

Query correctness, variables, units, data-source availability, provisioning consistency.

6 lessons

Part LXXXIX

Observability Platform Monitoring Itself

Who monitors the monitoring system; Prometheus health, scrape failures, ingest, storage, collectors.

6 lessons

Part XC

Meta-Monitoring

Designing the strategy where the observability stack detects failures in itself without circular assumptions.

6 lessons

Part XCI

Backup Strategy

What should and should not be backed up; configuration, dashboards, secrets, object storage, TSDB.

6 lessons

Part XCII

Disaster Recovery

RPO / RTO for observability; full-stack DR scenarios and rebuild procedures.

6 lessons

Part XCIII

Upgrades

Release notes, compatibility, backup, test, canary, validate, production; the upgrade playbook.

6 lessons

Part XCIV

Prometheus Upgrades

TSDB / data / config compatibility; migration patterns; rollback.

6 lessons

Part XCV

Grafana Upgrades

Database migrations, plugin compatibility, dashboards, provisioning, configuration.

6 lessons

Part XCVI

Loki Upgrades

Component / storage / schema / config compatibility per current versions.

6 lessons

Part XCVII

Tempo Upgrades

Safe compatibility checks; storage, receivers, query, retention.

6 lessons

Part XCVIII

Troubleshooting Methodology

Symptom → impact → signal → evidence → hypothesis → query → correlate → validate → root cause.

6 lessons

Part XCIX

Missing Metrics

Exporter → network → scrape → relabelling → target → query; the diagnostic path.

6 lessons

Part C

Missing Logs

Application → file / journal → collector → pipeline → Loki → query.

6 lessons

Part CI

Missing Traces

Instrumentation → propagation → collector → exporter → Tempo → sampling.

6 lessons

Part CII

Slow Queries

Large ranges, high cardinality, expensive regex, poor aggregation, storage bottlenecks.

6 lessons

Part CIII

Alert Failure

Production is broken but no alert fires; telemetry, rule, threshold, evaluator, Alertmanager, routing.

6 lessons

Part CIV

False Positive Alert

Investigation and tuning; alert fatigue as a failure mode.

6 lessons

Part CV

Cardinality Incident

Prometheus memory balloons; new label with request ID; investigate, redesign, recover.

6 lessons

Part CVI

Log Ingestion Incident

Loki ingestion volume rises 20×; debug level enabled, loop, new high-volume service, stack traces.

6 lessons

Part CVII

Trace Volume Incident

Excessive tracing; tail-sampler misconfig; 100% sampling; cost-driven failure.

6 lessons

Part CVIII

Clock Skew

How clock problems affect logs, traces, event ordering; cross-reference Linux time sync.

6 lessons

Part CIX

Incident Investigation Workflows

Complete investigations: alert → service metrics → host metrics → logs → traces → dependency metrics.

6 lessons

Part CX

Observability During Major Incidents

Incident dashboards, change annotations, query discipline, preserving evidence.

6 lessons

Part CXI

Observability Anti-Patterns

Monitor-everything, alert-on-everything, unlimited retention, high-cardinality labels, secrets-in-logs.

6 lessons

Part CXII

Production Observability Operating Model

Application teams → platform / SRE → observability platform → security / operations; ownership.

6 lessons

Part CXIII

Documentation and Runbooks

Alerts should link to impact, dashboard, runbook, owner; documentation as part of telemetry.

6 lessons

Part CXIV

Final Production Reference Architecture

End-to-end production observability stack with every component, dependency, failure domain.

6 lessons

Part Labs

Hands-On Labs

30+ labs across the course.

0 lessons

Part Runbooks

Operational Runbooks

30+ runbooks for production operations.

0 lessons

Part Checklists

Production Checklists

16+ checklists for production readiness.

0 lessons

Part Breakfix

Break/Fix Scenarios

30+ diagnostic and recovery exercises.

0 lessons

Part Capstone

Capstone

Production observability estate with failure injection and recovery.

0 lessons

Part Final

Final Assessment

Theory and practical assessment of every production competency.

0 lessons

Verified against

  • Prometheusv2.55.x· released 2026· verified 2026-08-13
  • Alertmanagerv0.28.x· released 2026· verified 2026-08-13
  • node_exporterv1.8.x· released 2026· verified 2026-08-13
  • blackbox_exporterv0.26.x· released 2026· verified 2026-08-13
  • Grafanav11.x· released 2026· verified 2026-08-13
  • Lokiv3.x· released 2026· verified 2026-08-13
  • Tempovcurrent· released 2026· verified 2026-08-13
  • OpenTelemetry Collectorv0.110.x· released 2026· verified 2026-08-13
  • Grafana Alloyvcurrent· released 2026· verified 2026-08-13
  • Docker Enginev28.x· released 2026· verified 2026-08-13
  • Ubuntuv24.04 LTS· verified 2026-08-13
  • Debianv12 (Bookworm)· verified 2026-08-13
  • RHEL / Rocky / AlmaLinuxv9.x· verified 2026-08-13