Skip to main content
RunBook Academy

Linux · Curriculum

Curriculum

523 lessons across 85 parts. Lessons build on each other; later parts assume familiarity with earlier material.

Part I

Foundations

UNIX and Linux history, kernel vs userland, distributions, FHS, /proc /sys /dev /run /etc /var /usr /home /boot.

6 lessons
  1. 01Welcome to Linux for Production SysadminsCourse introduction · foundation · ~8 min
  2. 02History and ecosystem of UNIX and LinuxHistory · foundation · ~10 min
  3. 03Kernel vs userland — the layered modelArchitecture · foundation · ~8 min
  4. 04The Filesystem Hierarchy StandardFilesystem hierarchy · foundation · ~10 minLab
  5. 05/proc, /sys, /dev, /run — the kernel-exposed filesystemsKernel filesystems · foundation · ~8 min
  6. 06Linux distributions and support lifecyclesDistributions · foundation · ~6 min

Part II

Shell and Command-Line Operations

Bash fluency, quoting, expansion, redirection, pipelines, exit codes, core text tools, command composition.

10 lessons
  1. 01Shell anatomy — commands, options, arguments, exit codesShell fundamentals · foundation · ~8 min
  2. 02Quoting, expansion, and globbingShell expansion · foundation · ~12 min
  3. 03Environment variables, PATH, and shell initialisationEnvironment · foundation · ~10 min
  4. 04Redirection, pipes, and three-channel I/OPipelines and redirection · foundation · ~12 minLab
  5. 05Exit codes, conditionals, and testsControl flow · foundation · ~10 min
  6. 06Core text tools — grep, cut, sort, uniq, tr, wcText tools · foundation · ~12 minLab
  7. 07sed and awk in productionText tools · intermediate · ~12 min
  8. 08find, locate, and xargs — bulk operations on filesBulk operations · intermediate · ~10 minLab
  9. 09Shell history, aliases, and functionsProductivity · foundation · ~8 min
  10. 10jq and structured text — JSON, YAML, and CSV at the command lineStructured text · intermediate · ~8 min

Part III

Filesystems and Files

Inodes, hard/symbolic links, permissions, umask, ACLs, extended attributes, immutable attributes, timestamps.

6 lessons
  1. 01Linux file types and the seven file kindsFile types · foundation · ~8 min
  2. 02Inodes, dentries, and the VFS — how Linux tracks filesInodes · foundation · ~12 min
  3. 03Hard links and symbolic linksLinks · foundation · ~10 min
  4. 04Permissions, ownership, and umaskPermissions · foundation · ~12 min
  5. 05POSIX ACLs — fine-grained permissions beyond owner/group/otherPermissions · intermediate · ~12 min
  6. 06Extended attributes, immutable attributes, and capabilitiesAttributes · intermediate · ~10 min

Part IV

Users, Groups and Identity

/etc/passwd /etc/shadow /etc/group, UIDs, GIDs, supplementary groups, password policies, service accounts.

6 lessons
  1. 01/etc/passwd, /etc/shadow, /etc/groupIdentity databases · foundation · ~10 min
  2. 02UIDs, GIDs, and supplementary groupsIdentity · foundation · ~10 min
  3. 03Password policies and account lockingIdentity · intermediate · ~10 min
  4. 04Service accounts and system accountsIdentity · intermediate · ~10 min
  5. 05Name Service Switch (NSS) and getentIdentity · foundation · ~8 min
  6. 06Account lifecycle and offboarding - creating, changing and removing usersIdentity · intermediate · ~16 min

Part V

sudo and Privileged Access

Root vs sudo, sudoers, /etc/sudoers.d, least privilege, logging, privileged session management.

6 lessons
  1. 01Root vs sudo — the production privilege modelPrivileged access · foundation · ~8 min
  2. 02sudoers syntax — rules, aliases, and defaultssudoers · intermediate · ~12 min
  3. 03Drop-in files, package-installed rules, and #includedirsudoers · intermediate · ~8 min
  4. 04Least-privilege sudo designsudo design · intermediate · ~12 min
  5. 05Sudo logging, I/O capture, and SIEM integrationsudo logging · intermediate · ~8 min
  6. 06Shell escapes, NOEXEC, and sudoeditsudo design · intermediate · ~14 min

Part VI

Processes

PID/PPID, sessions, process groups, threads, states, zombies, signals, priorities, nice values, cgroup views.

6 lessons
  1. 01The Linux process model — PIDs, PPIDs, sessions, and groupsProcess model · foundation · ~10 min
  2. 02Process lifecycle — fork, exec, exit, and waitProcess lifecycle · intermediate · ~12 min
  3. 03Signals — SIGTERM, SIGKILL, SIGHUP, and the restSignals · intermediate · ~12 min
  4. 04Threads and cgroup viewsThreads · intermediate · ~8 min
  5. 05Nice, renice, and CPU scheduling prioritiesPriorities · foundation · ~10 min
  6. 06ps, pgrep, pkill, top, htop — the production toolkitObservation tools · foundation · ~10 min

Part VII

systemd and Service Management

PID 1, units, services, targets, timers, sockets, mounts, dependencies, drop-in overrides, resource controls.

8 lessons
  1. 01systemd architecture — PID 1 and the unit modelArchitecture · intermediate · ~12 min
  2. 02systemctl and unit managementUnit management · foundation · ~10 min
  3. 03Writing systemd unit filesUnit authoring · intermediate · ~14 min
  4. 04Targets, dependencies, and orderingTransactions · intermediate · ~10 min
  5. 05systemd timers — replacing cron with calendar and monotonic schedulesTimers · intermediate · ~10 min
  6. 06systemd sockets, paths, and activationActivation · intermediate · ~10 min
  7. 07Drop-in overrides — modifying units without rewritingDrop-ins · intermediate · ~8 min
  8. 08systemd-analyze, security review, and boot diagnosticsDiagnostics · intermediate · ~10 min

Part VIII

Logging and journald

Kernel logs, journald, syslog, rsyslog, logrotate, persistent journals, filtering, central logging.

6 lessons
  1. 01Linux logging architecture — kernel, journald, syslogLogging architecture · foundation · ~10 min
  2. 02journalctl — querying the systemd journalJournal queries · foundation · ~12 min
  3. 03rsyslog, syslog-ng, and traditional syslogrsyslog · intermediate · ~10 min
  4. 04logrotate — the rotation policylogrotate · intermediate · ~10 min
  5. 05Persistent journals and disk pressureJournal sizing · intermediate · ~8 min
  6. 06Centralised logging — rsyslog forwarding and journald remoteCentral logging · intermediate · ~12 min

Part IX

Boot Process

Firmware, UEFI, GRUB, initramfs, kernel parameters, emergency and rescue targets, recovery.

6 lessons
  1. 01Firmware, UEFI, and BIOSPre-boot · foundation · ~10 min
  2. 02GRUB — the bootloaderBootloader · intermediate · ~12 min
  3. 03Kernel, initramfs, and the first userlandKernel + initramfs · intermediate · ~12 min
  4. 04systemd targets and rescue modeTargets · intermediate · ~10 min
  5. 05Emergency and rescue mode — when normal boot failsEmergency · intermediate · ~12 min
  6. 06Boot failure diagnosis from the OOB consoleDiagnosis · intermediate · ~10 min

Part X

Kernel Management

Kernel versions, modules, dependencies, command line, sysctl, kernel logs, taint, live patching concepts.

6 lessons
  1. 01Kernel versions, releases, and lifecycleKernel versions · foundation · ~10 min
  2. 02Kernel modules — load, list, dependencies, blacklistModules · intermediate · ~12 min
  3. 03sysctl — runtime kernel tuningsysctl · intermediate · ~10 min
  4. 04Kernel logs, taint flags, and oops decodingKernel logs · intermediate · ~10 min
  5. 05Live kernel patching — kpatch, kgraft, and livepatchLive patching · advanced · ~8 min
  6. 06Kernel command line parameters - the settings you can only make at bootKernel command line · intermediate · ~14 min

Part XI

Package Management

apt, dpkg, dnf, rpm, repositories, signing, dependency resolution, version pinning, rollback considerations.

7 lessons
  1. 01Package management overview — apt vs dnfPackage management · foundation · ~15 min
  2. 02apt and dpkg — Debian-family package managementapt and dpkg · foundation · ~12 min
  3. 03dnf and rpm — RHEL-family package managementdnf and rpm · foundation · ~12 min
  4. 04Repositories, signing, and trusted sourcesRepositories · intermediate · ~10 min
  5. 05Package pinning, holds, and version selectionVersion pinning · intermediate · ~10 min
  6. 06Security updates and CVE responseSecurity updates · intermediate · ~10 min
  7. 07Package cache, cleanup, and rollbackCache and rollback · foundation · ~8 min

Part XII

Repository Security and Supply Chain

GPG keys, package provenance, third-party repositories, malicious packages, dependency compromise.

6 lessons
  1. 01Repository trust and GPG signingTrust · foundation · ~10 min
  2. 02Third-party repository riskThird-party repos · intermediate · ~10 min
  3. 03Package provenance and SBOMProvenance · intermediate · ~10 min
  4. 04Mitigating curl | bash and other supply-chain footgunsSupply chain · intermediate · ~8 min
  5. 05Malicious packages and maintainer compromiseSupply chain · advanced · ~16 min
  6. 06Language package managers on production hostsSupply chain · advanced · ~16 min

Part XIII

Disks and Block Devices

Block devices, naming, sectors, partitions, GPT, signatures, UUIDs, labels, persistent naming.

6 lessons
  1. 01Block devices, naming, and disk identificationBlock devices · foundation · ~10 min
  2. 02Partitions, GPT, and MBRPartitioning · foundation · ~10 min
  3. 03Filesystem signatures and wipefsFilesystem signatures · foundation · ~8 min
  4. 04Persistent naming - UUIDs, labels, and stable pathsPersistent naming · foundation · ~8 min
  5. 05Sector sizes and partition alignment - 512n, 512e, and 4KnSectors and alignment · intermediate · ~15 min
  6. 06Rescanning and growing block devices - when the kernel still sees the old sizeRescan and resize · advanced · ~15 min

Part XIV

Filesystems

ext4, XFS, Btrfs characteristics, journaling, mount options, online growth, inode exhaustion, fsck/xfs_repair.

6 lessons
  1. 01Filesystem concepts - journaling, allocation, and durabilityFilesystem concepts · foundation · ~10 min
  2. 02ext4 deep dive - features, tuning, and operationsext4 · intermediate · ~12 min
  3. 03XFS deep dive - features, tuning, and operationsxfs · intermediate · ~12 min
  4. 04btrfs where it fits - snapshots, subvolumes, and operationsbtrfs · intermediate · ~10 min
  5. 05Inode exhaustion - the filesystem that says noOperations · intermediate · ~8 min
  6. 06Filesystem check, repair, and recovery with fsck and xfs_repairOperations · intermediate · ~10 min

Part XV

/etc/fstab and Mount Management

UUID, LABEL, mount options, systemd interaction, network filesystems, recovery from broken fstab.

6 lessons
  1. 01fstab syntax - fields, options, and the dump/pass columnsfstab syntax · foundation · ~10 min
  2. 02Mount options and systemd integrationMount options · intermediate · ~10 min
  3. 03Recovering from a broken fstab without a live USBBoot recovery · intermediate · ~10 min
  4. 04Network filesystems in fstab - _netdev, automount, and the boot you cannot watchNetwork filesystems · intermediate · ~15 min
  5. 05fstab entries that are not disks - bind mounts, tmpfs, and swapNon-device entries · intermediate · ~15 min
  6. 06Mount units by hand, and verifying a mount change before you rebootMount units · advanced · ~15 min

Part XVI

LVM

PV/VG/LV, extension, snapshots, risks of shrinking, recovery, day-2 operations.

7 lessons
  1. 01LVM concepts - PV, VG, LV, and the storage stackLVM concepts · foundation · ~10 min
  2. 02LVM creation, extension, and snapshotsLVM operations · intermediate · ~12 min
  3. 03LVM snapshots and rollback - the local safety netSnapshots · intermediate · ~8 min
  4. 04LVM monitoring and capacity planningMonitoring · intermediate · ~8 min
  5. 05LVM troubleshooting - common failures and how to recoverTroubleshooting · intermediate · ~8 min
  6. 06Shrinking an LV - the highest-risk LVM operationLVM operations · advanced · ~10 min
  7. 07LVM thin provisioning and thin-pool monitoringMonitoring · advanced · ~14 min

Part XVII

Software RAID

mdadm, RAID1/5/6/10, degraded arrays, rebuilds, failure detection, monitoring.

6 lessons
  1. 01mdadm RAID concepts - levels, layouts, and failure modesRAID concepts · foundation · ~10 min
  2. 02RAID1 mirroring - simple and reliableRAID levels · foundation · ~8 min
  3. 03RAID5 and RAID6 - parity for capacity and fault toleranceRAID levels · intermediate · ~10 min
  4. 04RAID10 - mirrored stripes for performance and redundancyRAID levels · intermediate · ~8 min
  5. 05Degraded arrays and rebuilds - replacing a failed diskRAID operations · intermediate · ~8 min
  6. 06md RAID monitoring and failure detectionRAID operations · intermediate · ~14 min

Part XVIII

Enterprise Storage

SAN, NAS, iSCSI, Fibre Channel, NFS, multipath, device-mapper, shared storage failure paths.

6 lessons
  1. 01SAN vs NAS vs object storage - which one to chooseStorage taxonomy · foundation · ~8 min
  2. 02iSCSI - block storage over IPiSCSI · intermediate · ~10 min
  3. 03Fibre Channel - high-speed block storageFC · intermediate · ~8 min
  4. 04NFS client - mounting and operating network filesystemsNFS · intermediate · ~10 min
  5. 05Multipath and device-mapper - redundant paths to SAN LUNsMultipath · intermediate · ~8 min
  6. 06Shared storage failure paths - reading a host whose storage has gone awayFailure paths · advanced · ~18 min

Part XIX

Networking Foundations

Ethernet, MAC, ARP, IPv4, IPv6, subnetting, routing, TCP, UDP, ICMP, DNS, DHCP concepts.

8 lessons
  1. 01OSI and TCP/IP models - what each layer is forModels · foundation · ~10 min
  2. 02Ethernet and MAC addressing - what is on the wireEthernet · foundation · ~10 min
  3. 03IPv4 addressing, subnetting, and CIDRIPv4 · foundation · ~14 min
  4. 04IPv6 addressing and subnettingIPv6 · foundation · ~12 min
  5. 05ARP and IPv6 neighbour discoveryARP/ND · foundation · ~10 min
  6. 06TCP and UDP - transport-layer essentialsTCP/UDP · foundation · ~14 min
  7. 07ICMP and path MTU - control messages and fragmentationICMP/MTU · foundation · ~10 min
  8. 08DNS and DHCP conceptsDNS/DHCP · foundation · ~12 min

Part XX

Linux Network Configuration

Interfaces, addresses, routes, gateways, DNS, Netplan, NetworkManager, systemd-networkd, persistent config.

7 lessons
  1. 01The ip command - the sysadmin network toolkitip command · foundation · ~12 min
  2. 02Linux interface configuration - addresses, MTU, offloadsInterface config · foundation · ~12 min
  3. 03Routing and default gateways on LinuxRouting · foundation · ~12 min
  4. 04Netplan on Ubuntu - declarative network configurationNetplan · foundation · ~10 min
  5. 05NetworkManager - flexible configuration for desktops and laptopsNetworkManager · foundation · ~10 min
  6. 06systemd-networkd - predictable configuration for serverssystemd-networkd · foundation · ~10 min
  7. 07Distribution comparison - network configuration across distrosComparison · foundation · ~8 min

Part XXI

Advanced Linux Networking

VLANs, bridges, bonding, teaming, MTU, jumbo frames, multiple routing tables, policy routing, IPv6, LACP.

8 lessons
  1. 01VLANs and 802.1Q - segmenting a network with tagsVLAN · intermediate · ~12 min
  2. 02Linux bridges and bridge VLANsBridge · intermediate · ~12 min
  3. 03Bonding and LACP - aggregating links for redundancy and throughputBond · intermediate · ~14 min
  4. 04NIC teaming - the userspace alternative to bondingTeaming · intermediate · ~8 min
  5. 05MTU and jumbo frames - when 1500 is not enoughMTU · intermediate · ~10 min
  6. 06Multiple routing tables on LinuxRouting tables · advanced · ~12 min
  7. 07Policy routing - choosing routes by source, mark, or interfacePolicy routing · advanced · ~12 min
  8. 08IPv6 advanced - privacy, SLAAC details, and operational gotchasIPv6 advanced · advanced · ~10 min

Part XXII

Network Troubleshooting

OSI/TCP-IP layer methodology, ip, ss, mtr, dig, tcpdump, ethtool, arping, nc, openssl s_client, evidence-based diagnosis.

9 lessons
  1. 01Network troubleshooting methodology - the systematic approachMethodology · foundation · ~14 min
  2. 02ss and socket state - inspecting TCP and UDP from the kernelss · foundation · ~12 min
  3. 03ping, traceroute, and mtr - path diagnosisPath diagnosis · foundation · ~12 min
  4. 04dig and nslookup - DNS diagnosis in detailDNS · foundation · ~12 min
  5. 05tcpdump and packet capture - seeing the bytestcpdump · intermediate · ~14 min
  6. 06ethtool and offloads - link-layer diagnosticsethtool · intermediate · ~12 min
  7. 07arping and ARP inspection - forcing layer-2 discoveryARP · intermediate · ~8 min
  8. 08nc and openssl s_client - probing TCP and TLSProbing · intermediate · ~12 min
  9. 09curl as a diagnostic tool - HTTP, headers, and timingcurl · foundation · ~10 min

Part XXIII

DNS

Resolver architecture, /etc/resolv.conf, systemd-resolved, recursive resolution, authoritative DNS, TTL, caching, cluster failure modes.

6 lessons
  1. 01DNS resolver architecture - how Linux finds namesArchitecture · foundation · ~12 min
  2. 02/etc/resolv.conf - the resolver configuration fileresolv.conf · foundation · ~10 min
  3. 03systemd-resolved - the modern stub resolversystemd-resolved · intermediate · ~15 min
  4. 04Recursive resolution - how a name becomes an IPResolution · foundation · ~10 min
  5. 05Authoritative DNS, TTLs, and caching behaviourAuthoritative · intermediate · ~12 min
  6. 06Cluster DNS failure modes - what breaks when DNS is downCluster DNS · advanced · ~10 min

Part XXIV

Time Synchronisation

UTC, monotonic time, NTP, chrony, systemd-timesyncd, operational impact on auth, TLS, logs, distributed systems.

6 lessons
  1. 01Time concepts - UTC, monotonic, and the wall clockConcepts · foundation · ~10 min
  2. 02chrony architecture - the modern NTP client and serverchrony · intermediate · ~12 min
  3. 03chrony client configuration - tuning for accuracy and resiliencechrony config · intermediate · ~12 min
  4. 04systemd-timesyncd - the lightweight NTP clientsystemd-timesyncd · intermediate · ~8 min
  5. 05Time skew operational impact - what breaks when clocks driftOperational impact · intermediate · ~10 min
  6. 06NTP fleet design - hierarchy, falsetickers and leap secondsFleet design · advanced · ~15 min

Part XXV

Firewalls

nftables as the modern foundation, plus iptables/firewalld/ufw differences, chains, hooks, stateful filtering, connection tracking, NAT.

9 lessons
  1. 01Firewall concepts - what a Linux firewall actually doesConcepts · foundation · ~10 min
  2. 02nftables architecture - tables, chains, and rulesnftables · intermediate · ~14 min
  3. 03nftables chains, tables, and hooks in depthnftables chains · intermediate · ~12 min
  4. 04nftables stateful filtering and conntracknftables stateful · intermediate · ~12 min
  5. 05nftables NAT - source NAT, destination NAT, and masqueradenftables NAT · intermediate · ~12 min
  6. 06iptables legacy and migration to nftablesiptables · intermediate · ~10 min
  7. 07firewalld - zone-based firewall managementfirewalld · intermediate · ~10 min
  8. 08ufw and distribution firewall wrappersufw · intermediate · ~8 min
  9. 09Validating actual firewall behaviourValidation · intermediate · ~10 min

Part XXVI

SSH

Client/server, keys, host keys, agent, forwarding, ProxyJump, configuration, certificates, MFA concepts, bastion hosts.

8 lessons
  1. 01SSH architecture - how Secure Shell actually worksArchitecture · foundation · ~10 min
  2. 02SSH keys and known_hosts - the key-based authentication modelKeys · foundation · ~12 min
  3. 03ssh-agent and forwarding - keeping keys safeAgent · intermediate · ~10 min
  4. 04sshd configuration - hardening the serversshd config · intermediate · ~14 min
  5. 05SSH certificates and bastionsCertificates · intermediate · ~12 min
  6. 06ProxyJump and jump hosts - secure access patternsProxyJump · intermediate · ~10 min
  7. 07SSH hardening - production checklistHardening · intermediate · ~12 min
  8. 08SSH MFA and 2FA conceptsMFA · intermediate · ~10 min

Part XXVII

Authentication and Enterprise Identity

PAM, NSS, LDAP, Active Directory, Kerberos, SSSD, central identity failure modes.

8 lessons
  1. 01PAM architecture - pluggable authentication on LinuxPAM · intermediate · ~12 min
  2. 02NSS - Name Service Switch for users, groups, and hostsNSS · intermediate · ~10 min
  3. 03LDAP and Active Directory - directory services fundamentalsLDAP · intermediate · ~12 min
  4. 04Kerberos concepts - tickets, realms, and principalsKerberos · advanced · ~12 min
  5. 05SSSD architecture - the system security services daemonSSSD · advanced · ~12 min
  6. 06SSSD and Active Directory integrationSSSD AD · advanced · ~14 min
  7. 07Central identity failure modes - what breaks when LDAP or AD is downIdentity failure · advanced · ~10 min
  8. 08PAM and NSS troubleshooting - fixing auth that breaksPAM NSS debug · intermediate · ~10 min

Part XXVIII

SELinux and AppArmor

Mandatory access control, contexts, policies, enforcement, audit logs, denial investigation, profile creation.

6 lessons
  1. 01MAC vs DAC - mandatory versus discretionary access controlConcepts · foundation · ~10 min
  2. 02SELinux architecture - contexts, policy, and enforcementSELinux · advanced · ~12 min
  3. 03SELinux contexts and labels - inspecting and fixingSELinux contexts · advanced · ~10 min
  4. 04SELinux policy and modes - enforcing, permissive, and custom policiesSELinux policy · advanced · ~12 min
  5. 05AppArmor concepts - path-based MAC for Debian-family systemsAppArmor · intermediate · ~10 min
  6. 06MAC denial investigation - diagnosing SELinux and AppArmor denialsMAC debug · intermediate · ~10 min

Part XXIX

Linux Security Hardening

Least privilege, SSH, sudo, mount options, kernel parameters, firewall, MAC, audit, packages, CIS-style controls.

7 lessons
  1. 01Linux hardening principles - defence in depthPrinciples · foundation · ~10 min
  2. 02SSH and sudo hardening - the privileged access surfaceSSH sudo · intermediate · ~12 min
  3. 03Kernel hardening with sysctl - the runtime network and security parameterssysctl · intermediate · ~10 min
  4. 04Mount options for security - nodev, nosuid, noexecMount options · foundation · ~8 min
  5. 05Service minimisation - reducing the attack surfaceService minimisation · intermediate · ~10 min
  6. 06Package update discipline - patching as a security controlUpdates · foundation · ~10 min
  7. 07CIS benchmarks conceptually - mapping controls to the frameworkCIS · intermediate · ~10 min

Part XXX

Linux Capabilities and Privilege

Traditional root model, Linux capabilities, file capabilities, process capabilities, capsh, getcap, setcap.

6 lessons
  1. 01Traditional root model - why capabilities existRoot model · foundation · ~10 min
  2. 02Linux capabilities overview - the full list and categoriesCapabilities overview · intermediate · ~12 min
  3. 03File capabilities - replacing setuid with fine-grained privilegesFile capabilities · intermediate · ~12 min
  4. 04capsh and capability debugging - inspecting and dropping capabilitiesCapability debug · intermediate · ~10 min
  5. 05Process capability sets and the exec transitionProcess capabilities · advanced · ~16 min
  6. 06Root-equivalent capabilities - when dropping to one changes nothingProcess capabilities · advanced · ~15 min

Part XXXI

Audit and Security Logging

auditd, authentication logs, sudo logs, kernel security events, ausearch, aureport, auditctl, central security logging.

6 lessons
  1. 01auditd architecture - the Linux audit frameworkauditd · intermediate · ~12 min
  2. 02audit rules and watch points - what to recordAudit rules · intermediate · ~12 min
  3. 03ausearch and aureport - querying the audit logAusearch aureport · foundation · ~10 min
  4. 04Authentication and sudo logs - what gets recorded and how to read itAuth logs · foundation · ~10 min
  5. 05Central security logging - shipping logs to a SIEMCentral logging · intermediate · ~10 min
  6. 06Audit record anatomy and attribution - answering whoAudit analysis · advanced · ~16 min

Part XXXII

Vulnerability and Patch Management

Detection, CVEs, severity vs context, exploitability, patch priority, maintenance windows, reboot requirements.

6 lessons
  1. 01Vulnerability concepts - CVEs, severity, and exploitabilityConcepts · foundation · ~10 min
  2. 02CVE severity vs context - prioritising for the environmentSeverity vs context · intermediate · ~10 min
  3. 03Scanning and detecting vulnerabilitiesScanning · intermediate · ~12 min
  4. 04Patch priority decisions - timing and riskPatch priority · intermediate · ~10 min
  5. 05Emergency patching - reacting to critical CVEs under pressureEmergency · advanced · ~10 min
  6. 06Reboot and restart requirements - finishing the patchPatch execution · intermediate · ~15 min

Part XXXIII

Fleet Patch Management

Dev/test/canary/production waves, staged deployments, health validation, rollback, automated patching trade-offs.

6 lessons
  1. 01Fleet patch strategy - waves, canaries, and rollbackStrategy · intermediate · ~10 min
  2. 02Canary and wave deployments - phased rollout with health checksCanary waves · intermediate · ~10 min
  3. 03Rollback strategies - safe deployment reversalsRollback · intermediate · ~10 min
  4. 04Pre-patching and post-patching validationValidation · intermediate · ~10 min
  5. 05Automated patching in a fleet - what to automate and what never toAutomation · advanced · ~16 min
  6. 06Repository snapshots - making a wave reproducibleAutomation · advanced · ~17 min

Part XXXIV

Configuration Management

Desired state, idempotency, drift, inventory, automation, Ansible-flavoured examples.

6 lessons
  1. 01Why manual does not scale - the case for configuration managementWhy CM · foundation · ~10 min
  2. 02Desired state and idempotency - the foundations of CMDesired state · intermediate · ~10 min
  3. 03Ansible for sysadmins - the modern CM toolAnsible · intermediate · ~12 min
  4. 04Inventory and facts - knowing your fleetInventory · intermediate · ~10 min
  5. 05Config drift detection - keeping reality aligned with declared stateDrift detection · intermediate · ~10 min
  6. 06Safe CM rollout - blast radius, check mode, and the control nodeSafe rollout · advanced · ~15 min

Part XXXV

Shell Scripting for Sysadmins

Variables, conditions, loops, functions, traps, exit codes, error handling, set -euo pipefail, ShellCheck.

7 lessons
  1. 01Script structure and conventions - the production shell scriptStructure · intermediate · ~10 min
  2. 02Variables and quoting - the foundation of safe shell scriptsVariables · intermediate · ~10 min
  3. 03Conditions and loops - controlling script flowConditions loops · foundation · ~10 min
  4. 04Functions and arguments - reusable, testable script componentsFunctions · intermediate · ~10 min
  5. 05Error handling - set -euo pipefail and trapError handling · intermediate · ~12 min
  6. 06Traps and temporary files - clean exit on signal or errorTraps tmpfiles · intermediate · ~10 min
  7. 07ShellCheck and static analysis - catching bugs before runtimeShellCheck · foundation · ~10 min

Part XXXVI

Scheduled Operations

cron, anacron, systemd timers, locking, duplicate prevention, logging, failure detection.

6 lessons
  1. 01cron syntax and anacron - the classic schedulercron · foundation · ~10 min
  2. 02systemd timers replacing cron - the modern schedulersystemd timers · intermediate · ~12 min
  3. 03Locking and concurrency - preventing duplicate jobsLocking · intermediate · ~10 min
  4. 04Job logging and failure detection - knowing what happenedLogging · intermediate · ~10 min
  5. 05Timeouts, retries and jitter - scheduled jobs that fail safelyJob safety · intermediate · ~18 min
  6. 06Auditing the schedule - what actually runs on this hostSchedule inventory · intermediate · ~17 min

Part XXXVII

Resource Management

CPU, memory, swap, processes, file descriptors, ulimits, cgroups, systemd resource controls, OOM behaviour.

6 lessons
  1. 01cgroups v2 architecture - the Linux resource control subsystemcgroups · advanced · ~12 min
  2. 02ulimits and limits.conf - per-user resource limitsulimits · foundation · ~10 min
  3. 03File descriptors and /proc/fd - understanding open filesFile descriptors · intermediate · ~14 min
  4. 04systemd resource controls - applying cgroups via unitssystemd resource controls · intermediate · ~12 min
  5. 05OOM killer behaviour - what happens when memory runs outOOM killer · advanced · ~10 min
  6. 06cgroup accounting - proving which limit was actually hitResource accounting · advanced · ~20 min

Part XXXVIII

Linux Performance Fundamentals

USE methodology, top, htop, vmstat, mpstat, pidstat, iostat, sar, free, slabtop, evidence-based investigation.

6 lessons
  1. 01USE methodology - the framework for performance investigationUSE method · foundation · ~10 min
  2. 02top, htop, and vmstat - the basic performance toolstop htop vmstat · foundation · ~10 min
  3. 03mpstat and pidstat - per-CPU and per-process statisticsmpstat pidstat · intermediate · ~10 min
  4. 04iostat and sar - storage and historical performanceiostat sar · intermediate · ~10 min
  5. 05free, slabtop, and /proc/meminfo - memory diagnosticsfree slabtop meminfo · intermediate · ~10 min
  6. 06Measure, change one thing, measure again - the tuning disciplineMethod · intermediate · ~19 min

Part XXXIX

CPU Performance

Utilisation, load average, run queue, context switching, interrupts, softirqs, steal time, CPU affinity, NUMA.

6 lessons
  1. 01CPU utilisation and load average - measuring the workloadCPU utilisation · foundation · ~10 min
  2. 02Run queue and context switching - measuring contentionRun queue · intermediate · ~10 min
  3. 03Interrupts, softirqs, and steal time - hardware and VM signalsInterrupts steal · intermediate · ~10 min
  4. 04CPU affinity and NUMA - controlling where processes runCPU affinity · intermediate · ~10 min
  5. 05CPU frequency, governors and thermal throttling - the clock is not constantFrequency · advanced · ~19 min
  6. 06perf - finding the code that is burning the CPUProfiling · advanced · ~21 min

Part XL

Memory Performance

Virtual memory, pages, page cache, anonymous memory, buffers, swap, memory pressure, OOM killer, slab.

6 lessons
  1. 01Virtual memory and pages - how Linux manages memoryVirtual memory · intermediate · ~10 min
  2. 02Page cache and anonymous memory - the two memory kindsPage cache anon · intermediate · ~10 min
  3. 03Swap and memory pressure - when the system runs out of RAMSwap · intermediate · ~10 min
  4. 04OOM killer decisions - who gets killed and whyOOM killer decisions · advanced · ~10 min
  5. 05RSS, PSS and shared pages - what a process really costsProcess footprint · advanced · ~19 min
  6. 06Kernel memory, huge pages and the tuning knobs that backfireKernel memory · advanced · ~21 min

Part XLI

Storage Performance

Latency, throughput, IOPS, queue depth, utilisation, filesystem cache, fio, iostat, iotop, safe benchmarking.

6 lessons
  1. 01I/O latency, throughput, and IOPS - the three storage metricsI/O metrics · foundation · ~10 min
  2. 02Queue depth and utilisation - storage saturation signalsQueue depth · intermediate · ~16 min
  3. 03iotop and pidstat - per-process I/O statisticsiotop pidstat · intermediate · ~10 min
  4. 04fio and safe benchmarking - measuring storage baselinefio · intermediate · ~12 min
  5. 05The I/O stack - which layer is adding the latencyI/O stack · advanced · ~21 min
  6. 06Writeback and filesystem tuning - dirty pages, mount options and fstrimFilesystem tuning · advanced · ~20 min

Part XLII

Network Performance

Bandwidth, latency, packet loss, retransmissions, socket queues, connection states, iperf3, ethtool, sar.

6 lessons
  1. 01Bandwidth and latency - the two network metricsBandwidth latency · foundation · ~10 min
  2. 02Packet loss and retransmits - when the network dropsPacket loss · intermediate · ~10 min
  3. 03Socket queues and connection states - TCP tuningSocket queues · intermediate · ~10 min
  4. 04iperf3 and throughput testing - measuring bandwidthiperf3 · foundation · ~10 min
  5. 05The bandwidth-delay product - why a fast link runs slowLong fat networks · advanced · ~20 min
  6. 06sar -n - network history and the counters nobody collectedHistorical data · advanced · ~18 min

Part XLIII

eBPF and Advanced Observability

eBPF, tracepoints, kprobes, uprobes, BPF maps, bpftrace, BCC, sysadmin use cases.

6 lessons
  1. 01eBPF overview - in-kernel observability for sysadminseBPF · advanced · ~12 min
  2. 02Tracepoints, kprobes, and uprobes - eBPF attachment pointsTracepoints kprobes · advanced · ~10 min
  3. 03bpftrace and bcc - sysadmin use cases and examplesbpftrace bcc · advanced · ~10 min
  4. 04BPF maps and state - communicating with userspaceBPF maps · advanced · ~10 min
  5. 05Running eBPF safely - kernel requirements, permissions and lockdownRequirements · expert · ~21 min
  6. 06The cost of tracing - eBPF overhead, dropped events and production safetyProduction safety · expert · ~21 min

Part XLIV

Central Monitoring

Prometheus node_exporter, Grafana, alerting on CPU, memory, FS, inodes, disk latency, network, services, time sync, hardware.

6 lessons
  1. 01Monitoring design - what to monitor, how to alert, and whyMonitoring design · foundation · ~10 min
  2. 02Prometheus and node_exporter - the standard Linux monitoring stackPrometheus · intermediate · ~10 min
  3. 03Grafana and alerting - dashboards and notificationGrafana alerting · intermediate · ~10 min
  4. 04What to monitor on Linux - the production checklistWhat to monitor · foundation · ~10 min
  5. 05Hardware monitoring sensors - temperature, voltage, fanHardware sensors · foundation · ~10 min
  6. 06Operating the monitoring stack - retention, cardinality and who watches the watcherMonitoring design · advanced · ~15 min

Part XLV

Central Logging

journald, rsyslog, syslog, Fluent Bit, Vector, Loki, Elasticsearch, retention, filtering, cardinality, capacity.

6 lessons
  1. 01Central logging design - the architecture and decisionsLogging design · foundation · ~10 min
  2. 02rsyslog forwarding - traditional syslog shippingrsyslog · intermediate · ~10 min
  3. 03Fluent Bit and Vector - modern log shippersFluent Bit Vector · intermediate · ~10 min
  4. 04Loki and Elasticsearch - the modern log storeLoki Elasticsearch · intermediate · ~10 min
  5. 05Log cardinality and capacity - sizing and tuningCardinality capacity · advanced · ~10 min
  6. 06Where log messages go missing - and how to prove they did notCentral logging · advanced · ~15 min

Part XLVI

OpenTelemetry

Metrics, logs, traces, collectors, agents/gateways, Linux infrastructure telemetry.

6 lessons
  1. 01OpenTelemetry three pillars - metrics, logs, tracesOTel three pillars · foundation · ~10 min
  2. 02OTel collectors and agents - the deployment modelOTel collectors · intermediate · ~10 min
  3. 03OTel host receiver and process metricsOTel host receiver · intermediate · ~10 min
  4. 04OTel into Grafana and Tempo - the visual stackOTel visualisation · intermediate · ~10 min
  5. 05OTel logs on a Linux host - the journald and filelog receiversOTel collectors · advanced · ~14 min
  6. 06The cost of telemetry - cardinality, sampling and collector limitsOTel collectors · advanced · ~15 min

Part XLVII

Backup Strategy

Filesystem, application, database, configuration, snapshot, consistency, encryption, retention, immutable copies, offsite, 3-2-1.

6 lessons
  1. 01Backup concepts - what to back up and howConcepts · foundation · ~10 min
  2. 02Application-consistent vs crash-consistent backupsConsistency · intermediate · ~10 min
  3. 033-2-1 and modern backup strategies3-2-1 · foundation · ~10 min
  4. 04Immutable and offline copies - the last line of defenceImmutable offline · intermediate · ~10 min
  5. 05Encryption and key management - protecting backups at restEncryption keys · advanced · ~10 min
  6. 06Backing up what is not a file - partition tables, LVM metadata, LUKS headers and package stateStructural metadata · advanced · ~18 min

Part XLVIII

Backup Tools

rsync, tar, Borg, Restic, filesystem snapshots, enterprise backup integration, selection criteria.

6 lessons
  1. 01rsync and tar - the classic backup toolsrsync tar · foundation · ~10 min
  2. 02BorgBackup and Restic - the modern deduplicated backup toolsBorg Restic · intermediate · ~12 min
  3. 03Filesystem snapshots as backup - ZFS and btrfsFS snapshots · intermediate · ~10 min
  4. 04LVM snapshots for backups - the classic Linux approachLVM snapshots · intermediate · ~10 min
  5. 05Enterprise backup integration - Veeam, NetBackup, and similarEnterprise · intermediate · ~10 min
  6. 06Backup tool selection - choosing the right tool for the workloadTool selection · intermediate · ~10 min

Part XLIX

Restore

Restoration over backup success: files, ownership, ACLs, services, configuration, full restore exercises.

6 lessons
  1. 01Restore matters more than backup - the discipline of restoreRestore matters · foundation · ~10 min
  2. 02Restore testing discipline - the regular validation routineRestore testing · intermediate · ~10 min
  3. 03Restore files and ownership - the practical detailsFiles ownership · intermediate · ~10 min
  4. 04Restore services and configuration - the full-system recoveryServices config · advanced · ~10 min
  5. 05Partial and point-in-time restores - one file, one directory, one momentPartial restore · advanced · ~18 min
  6. 06Restore drills - turning an asserted RTO into a measured oneDrills and measurement · advanced · ~18 min

Part L

Disaster Recovery

RPO, RTO, disaster scenarios, rebuild vs restore, bare-metal recovery, DNS, certificates, identity, full cluster loss.

6 lessons
  1. 01RPO and RTO - modelling recovery objectivesRPO RTO · foundation · ~10 min
  2. 02Disaster scenarios - what to plan forDisaster scenarios · foundation · ~10 min
  3. 03Bare-metal recovery - restoring to a fresh hostBare metal · advanced · ~10 min
  4. 04Dependencies: DNS, certificates, identity - the cascade risksDependencies · advanced · ~10 min
  5. 05Full cluster loss recovery - the worst-case scenarioFull cluster loss · advanced · ~10 min
  6. 06Declaring a disaster, failing over, and failing backDeclaration and failback · expert · ~20 min

Part LI

Linux Fleet Architecture

Management plane, configuration management, monitoring, logging, identity, automation, secrets, patching across 50-5000 nodes.

6 lessons
  1. 01Fleet management plane - how a Linux fleet is operatedManagement plane · foundation · ~10 min
  2. 02Naming and inventory - knowing your fleetNaming inventory · foundation · ~10 min
  3. 03Environment separation - dev, staging, and productionEnvironment separation · foundation · ~10 min
  4. 04Central identity and secrets - the foundation of fleet securityCentral identity secrets · advanced · ~10 min
  5. 05Host enrolment and decommissioning - joining and leaving the fleetFleet lifecycle · advanced · ~16 min
  6. 06Scaling the management plane - what breaks between 50 and 5000 nodesFleet lifecycle · expert · ~17 min

Part LII

High Availability Fundamentals

Availability, redundancy, fault tolerance, failure domains, active/active, active/passive, N+1, N+2, host vs service availability.

6 lessons
  1. 01Availability and reliability - the metrics that matterAvailability · foundation · ~10 min
  2. 02Redundancy and fault tolerance - building for failureRedundancy · foundation · ~10 min
  3. 03Failure domains - what fails togetherFailure domains · intermediate · ~10 min
  4. 04N+1 capacity - the math of headroomN+1 · intermediate · ~10 min
  5. 05Host availability vs service availability - measuring the right thingHost vs service · intermediate · ~12 min
  6. 06Redundancy levels - N+1, N+2, 2N and 2N+1Redundancy levels · intermediate · ~12 min

Part LIII

Quorum and Split Brain

Quorum, majority, membership, partitions, split brain, witness/quorum device, failure detection.

6 lessons
  1. 01Quorum concepts - the math of agreementQuorum · advanced · ~10 min
  2. 02Membership and partitions - the failure mode that breaks quorumMembership · advanced · ~10 min
  3. 03Split brain explained - the cluster failure modeSplit brain · advanced · ~10 min
  4. 04Witness and quorum devices - breaking geographic tiesWitness · advanced · ~10 min
  5. 05Failure detection - how a cluster decides a node has failedFailure detection · advanced · ~13 min
  6. 06Two-node clusters and tie-breaking - the honest versionTwo-node clusters · advanced · ~14 min

Part LIV

Fencing and STONITH

Why fencing exists, split-brain data corruption, fencing devices, STONITH, hardware-independent principles.

6 lessons
  1. 01Why fencing exists - the cluster disciplineFencing · advanced · ~10 min
  2. 02Fencing devices and agents - the practical implementationFencing devices · advanced · ~10 min
  3. 03STONITH and data integrity - the shoot-the-other-node-in-the-head patternSTONITH · advanced · ~10 min
  4. 04Unresponsive is not dead - the evidence problem fencing solvesEvidence · advanced · ~14 min
  5. 05SBD and watchdog fencing - when the node fences itselfSBD · advanced · ~14 min
  6. 06Fencing design principles - evaluating any platformDesign principles · advanced · ~13 min

Part LV

Pacemaker and Corosync

Corosync membership, Pacemaker, resources, resource agents, constraints, failover, fencing integration.

6 lessons
  1. 01Corosync architecture - the cluster membership layerCorosync · advanced · ~10 min
  2. 02Pacemaker architecture - the resource managerPacemaker · advanced · ~10 min
  3. 03Pacemaker resources and resource agentsPacemaker resources · advanced · ~10 min
  4. 04Pacemaker constraints: location, colocation, orderPacemaker constraints · advanced · ~10 min
  5. 05Pacemaker fencing integration - tying STONITH to resourcesPacemaker fencing · advanced · ~10 min
  6. 06Pacemaker troubleshooting - the systematic approachTroubleshooting · advanced · ~10 min

Part LVI

Keepalived and VRRP

Virtual IPs, VRRP, health checks, master/backup, failover, limitations.

6 lessons
  1. 01VRRP concepts - virtual router redundancy protocolVRRP · foundation · ~10 min
  2. 02Keepalived configuration - the Linux VRRP implementationKeepalived config · intermediate · ~10 min
  3. 03Keepalived health checks - service-aware failoverHealth checks · intermediate · ~10 min
  4. 04Keepalived limitations - when VRRP is not enoughLimitations · intermediate · ~10 min
  5. 05Virtual IP mechanics - what actually moves during a failoverVIP mechanics · intermediate · ~14 min
  6. 06Sync groups and multiple VIPs - keeping a service togetherSync groups · intermediate · ~13 min

Part LVII

Linux Load Balancing

HAProxy, nginx, IPVS, Layer 4 vs Layer 7, health checking, algorithms, session persistence, connection draining.

6 lessons
  1. 01Load balancing concepts - distributing traffic across backendsConcepts · foundation · ~10 min
  2. 02HAProxy configuration - the production-grade layer 7 LBHAProxy · intermediate · ~12 min
  3. 03nginx as a load balancer - the versatile optionnginx LB · intermediate · ~10 min
  4. 04IPVS and the kernel load balancer - the layer 4 workhorseIPVS · advanced · ~10 min
  5. 05Health checks and persistence - the safety net of load balancingHealth checks · intermediate · ~10 min
  6. 06Connection draining - taking a backend out without dropping requestsLoad balancing · advanced · ~14 min

Part LVIII

Clustered Service Architecture

Stateless, shared state, replicated state, external state, application architecture constraints, identifying the right pattern.

6 lessons
  1. 01Stateless services - the simplest cluster architectureStateless · foundation · ~10 min
  2. 02Shared state and replicated state - when stateless is not enoughShared state · advanced · ~10 min
  3. 03External state and databases - the stateful backendExternal state · advanced · ~10 min
  4. 04Application architecture constraints - the limits of clusteringConstraints · advanced · ~10 min
  5. 05Identifying the right cluster pattern - a decision procedureChoosing a pattern · advanced · ~13 min
  6. 06Failover as the client sees it - the outage the cluster does not measureClient-visible recovery · advanced · ~13 min

Part LIX

Shared Storage and Clusters

NFS, SAN, clustered filesystems, distributed storage, two-nodes-mounting-the-same-block-filesystem risk.

6 lessons
  1. 01NFS as cluster storage - the simple shared filesystemNFS · intermediate · ~10 min
  2. 02SAN and cluster storage - the enterprise shared filesystemSAN · advanced · ~14 min
  3. 03Clustered filesystems concepts - POSIX over multiple hostsClustered FS · advanced · ~10 min
  4. 04Two nodes mounting the same block device - the data corruption riskTwo nodes risk · advanced · ~10 min
  5. 05Cluster-managed mounts - why shared storage never goes in fstabCluster mounts · advanced · ~13 min
  6. 06Shared storage failure modes - the single point of failure you boughtShared storage failure · advanced · ~14 min

Part LX

Distributed Storage Concepts

Ceph, Gluster, object storage, replicated block, replication, quorum, failure domains, consistency, recovery.

6 lessons
  1. 01Ceph architecture intro - the modern distributed storageCeph · advanced · ~10 min
  2. 02Gluster concepts intro - the other distributed filesystemGluster · intermediate · ~10 min
  3. 03Object and block storage tradeoffs - the right tool for the dataObject vs block · intermediate · ~10 min
  4. 04Replication, quorum, and consistency - the distributed storage disciplineReplication · advanced · ~10 min
  5. 05Failure domains in distributed storage - where the replicas actually landFailure domains · advanced · ~14 min
  6. 06Recovery and rebalancing - what a distributed store does after a failureRecovery · advanced · ~15 min

Part LXI

DRBD Concepts

Primary/secondary, replication, split-brain, clustering integration, when DRBD materially helps HA.

6 lessons
  1. 01DRBD architecture - the distributed replicated block deviceDRBD · advanced · ~10 min
  2. 02DRBD primary/secondary and split-brain handlingDRBD roles · advanced · ~16 min
  3. 03DRBD in a cluster - integration with Pacemaker for HADRBD cluster · advanced · ~14 min
  4. 04DRBD resync and online verification - reading the replication stateDRBD operations · advanced · ~14 min
  5. 05Dual-primary DRBD - what it actually requiresDRBD dual-primary · expert · ~14 min
  6. 06When DRBD materially helps HA - and when it does notDRBD decisions · advanced · ~12 min

Part LXII

Cluster Networking

Management, application, storage, heartbeat, OOB, redundant interfaces, switches, VLANs, failure domains.

6 lessons
  1. 01Cluster network design - management, application, and storageCluster network design · advanced · ~10 min
  2. 02Management vs application vs storage - the network rolesNetwork roles · intermediate · ~10 min
  3. 03Redundant interfaces and switches - the cluster network resilienceRedundancy · advanced · ~10 min
  4. 04VLANs and failure domains in cluster networksVLANs · intermediate · ~10 min
  5. 05Multicast and unicast on the cluster networkTransport · advanced · ~13 min
  6. 06Latency, jitter and cluster membershipTiming · expert · ~14 min

Part LXIII

Cluster Time, DNS and Identity Dependencies

DNS unavailable, NTP skew, LDAP unavailable, certificate expiry, designing around dependency failures.

6 lessons
  1. 01DNS cluster failure impact - what happens when DNS is downDNS failure · intermediate · ~10 min
  2. 02NTP skew cluster impact - the silent time bombNTP skew · intermediate · ~10 min
  3. 03LDAP and certificate expiry cluster impact - the silent credential failuresIdentity expiry · intermediate · ~10 min
  4. 04Mapping cluster dependencies - finding what you depend onDependency mapping · advanced · ~13 min
  5. 05Designing around dependency failures - remove, cache, degrade, break glassDesign · advanced · ~14 min
  6. 06Testing dependency failures - injecting the outage safelyTesting · advanced · ~14 min

Part LXIV

Rolling Maintenance

Validate, drain, patch, reboot, validate, return, observe, batch sizing, maintenance mode, health checks.

6 lessons
  1. 00Rolling maintenance workflow - the loop that keeps a cluster servingWorkflow · advanced · ~12 min
  2. 01Maintenance mode and drain - the safe cluster updateMaintenance mode · intermediate · ~10 min
  3. 02Health validation after a change - the safety netHealth validation · intermediate · ~10 min
  4. 03Batch sizing and rollback - the change control disciplineBatch and rollback · intermediate · ~10 min
  5. 04Draining behind a load balancer - rolling maintenance without a cluster managerDraining · advanced · ~14 min
  6. 05Orchestrating the rolling loop - automation that stops on evidenceOrchestration · advanced · ~15 min

Part LXV

Rolling Kernel Upgrades

Kernel package install, reboot requirement, boot validation, rollback, cluster capacity during maintenance, live patching concepts.

6 lessons
  1. 01Kernel package install - the kernel upgrade workflowKernel install · intermediate · ~10 min
  2. 02Reboot and boot validation - the kernel upgrade verificationBoot validation · intermediate · ~10 min
  3. 03Live kernel patching concepts - patching without rebootingLive patching · intermediate · ~10 min
  4. 04Kernel rollback - getting back to the kernel that workedRollback · advanced · ~12 min
  5. 05Cluster capacity during a kernel campaign - the reboot budgetCapacity · advanced · ~14 min
  6. 06Proving a kernel campaign worked - CVEs, backports and vulnerability stateVerification · advanced · ~14 min

Part LXVI

Capacity Planning for Clusters

Normal/peak utilisation, failure capacity, maintenance capacity, growth, N+1 headroom, 90%-on-3-nodes trap.

6 lessons
  1. 01Capacity utilisation bands - the 90%-on-3-nodes trapUtilisation bands · intermediate · ~10 min
  2. 02Normal and peak - measuring the load you actually haveMeasurement · intermediate · ~14 min
  3. 03Utilisation is not saturation - what the queue tells you that the percentage cannotMeasurement · advanced · ~15 min
  4. 04Forecasting growth - turning a trend into a dateForecasting · intermediate · ~15 min
  5. 05Maintenance capacity - what a change window costs the clusterFailure and maintenance · advanced · ~14 min
  6. 06Writing the capacity plan - finding the binding constraintThe plan · advanced · ~15 min

Part LXVII

Cluster Monitoring

Node health, service health, quorum, membership, failovers, resource status, storage, network, replication, time sync.

6 lessons
  1. 01Cluster monitoring design - what to monitor in a clusterDesign · advanced · ~10 min
  2. 02Monitoring failovers and resources - the cluster activityFailovers and resources · intermediate · ~14 min
  3. 03Actionable cluster alerts - the alert that leads to actionActionable alerts · intermediate · ~10 min
  4. 04Monitoring cluster storage and replication - watching the standby tooStorage and replication · intermediate · ~13 min
  5. 05Monitoring cluster network and time - the leading indicatorsNetwork and time · advanced · ~13 min
  6. 06What cluster monitoring cannot tell youLimits · advanced · ~13 min

Part LXVIII

Cluster Incident Response

Ten canonical scenarios: node unreachable, partition, quorum loss, fencing failure, shared storage loss, LB failure, clock skew, DNS outage, deployment regression, memory exhaustion.

6 lessons
  1. 01Cluster IR: node unreachable - the most common incidentNode unreachable · intermediate · ~10 min
  2. 02Cluster IR: network partition - deciding which side is realPartition · advanced · ~14 min
  3. 03Cluster IR: quorum loss - the arithmetic and the decision recordQuorum loss · advanced · ~14 min
  4. 04Cluster IR: fencing failure and the fence loopFencing failure · advanced · ~15 min
  5. 05Cluster IR: shared storage loss - when a stop cannot succeedShared storage loss · advanced · ~15 min
  6. 06Cluster IR: building one timeline from nodes whose clocks disagreeCross-node evidence · advanced · ~15 min

Part LXIX

Hardware Health

SMART, NVMe health, RAID controllers, IPMI, BMC, Redfish concepts, ECC memory, thermal sensors, smartctl, nvme, sensors, ipmitool.

6 lessons
  1. 01SMART and NVMe health - disk failure predictionSMART NVMe · intermediate · ~10 min
  2. 02RAID controller monitoring - the storage layer belowRAID controller · advanced · ~10 min
  3. 03IPMI and BMC basics - the lights-out management foundationIPMI BMC · foundation · ~10 min
  4. 04Thermal and ECC monitoring - the memory and CPU healthThermal ECC · intermediate · ~10 min
  5. 05Hardware inventory and the SEL - turning an alert into a part numberInventory and SEL · intermediate · ~13 min
  6. 06Firmware lifecycle - the layer your package manager cannot seeFirmware · advanced · ~14 min

Part LXX

Out-of-Band Management

IPMI, iDRAC, iLO, Redfish, serial console, remote console, power control, relation to fencing and DR.

6 lessons
  1. 01OOB architecture - the out-of-band management layerOOB architecture · foundation · ~10 min
  2. 02Redfish and IPMI APIs - the standards for hardware managementRedfish IPMI · advanced · ~10 min
  3. 03OOB and DR - the out-of-band layer in disaster recoveryOOB and DR · advanced · ~10 min
  4. 04Serial console and SOL - the text lifeline into a hostSerial console · advanced · ~14 min
  5. 05OOB power control - the commands that change statePower control · advanced · ~14 min
  6. 06OOB credentials and break-glass - access that survives the outageBreak-glass access · advanced · ~14 min

Part LXXI

TLS and PKI

Private keys, certificates, CSRs, CAs, chains, SAN, expiry, revocation, openssl, certificate troubleshooting, expiry incidents.

6 lessons
  1. 01TLS and PKI concepts - the foundation of encrypted communicationTLS and PKI · foundation · ~13 min
  2. 02OpenSSL for sysadmins - the practical toolkitOpenSSL · intermediate · ~14 min
  3. 03Certificate authorities and chains - the trust model in practiceCAs and chains · intermediate · ~16 min
  4. 04Certificate renewal and expiry - the lifecycle disciplineRenewal expiry · intermediate · ~12 min
  5. 05TLS troubleshooting - diagnosing failures from the wireTroubleshooting · advanced · ~17 min
  6. 06Certificate revocation in practice - CRL, OCSP and short lifetimesTroubleshooting · advanced · ~16 min

Part LXXII

Secrets

Passwords, private keys, API tokens, service credentials, secrets in scripts/Git/world-readable/shell history, external managers.

6 lessons
  1. 01Secrets - the credential management problemProblem · foundation · ~22 min
  2. 02Secrets at rest - ownership, modes and the copies you forgotAt rest · intermediate · ~16 min
  3. 03Secrets in scripts, logs and shell historyIn transit through your own tooling · intermediate · ~16 min
  4. 04Private keys - passphrases, agents and the keys you cannot rotateKey material · intermediate · ~18 min
  5. 05Rotation - making a credential change a routine operationLifecycle · intermediate · ~17 min
  6. 06External secret managers - the bootstrap and availability problemsExternal stores · advanced · ~18 min

Part LXXIII

Change Management

Change plans, peer review, pre-checks, rollback, maintenance windows, testing, validation, observation period.

6 lessons
  1. 01Change plans and peer review - the discipline of safe changesPlans and review · foundation · ~10 min
  2. 02Change classes and blast radius - sizing the process to the changeClassification · intermediate · ~14 min
  3. 03Pre-checks and state capture - what to record before you touch anythingExecution · intermediate · ~15 minLab
  4. 04Staged rollout - canaries, cohorts and the health gateExecution · advanced · ~15 minLab
  5. 05Abort criteria and the point of no return - rollback as part of the procedureExecution · advanced · ~15 min
  6. 06Maintenance windows, freezes and the observation periodScheduling · advanced · ~14 min

Part LXXIV

Configuration Drift

How environments become inconsistent, detection, remediation, drift across nodes, packages, firewall rules, configuration.

6 lessons
  1. 01Drift remediation - bringing hosts back to desired stateRemediation · intermediate · ~10 min
  2. 02How drift accumulates - the eight ways hosts stop matchingOrigins of drift · intermediate · ~13 min
  3. 03Detecting drift independently of the CM toolIndependent detection · intermediate · ~15 min
  4. 04Package and version drift across a fleetPackage drift · intermediate · ~14 min
  5. 05Runtime versus on-disk drift - when the file is right and the system is notRuntime drift · intermediate · ~15 min
  6. 06The limits of drift detection - measuring what nothing managesCoverage · advanced · ~15 min

Part LXXV

Immutable vs Mutable Infrastructure

Traditional servers, configuration management, golden images, cloud images, immutable replacement, trade-offs.

6 lessons
  1. 01Mutable server traditions - the legacy approachMutable server · foundation · ~10 min
  2. 02Golden images and immutable patterns - the modern approachGolden images · intermediate · ~10 min
  3. 03Cloud images and the image build pipelineImage pipeline · intermediate · ~17 min
  4. 04Configuration management and image baking - choosing the seamConfiguration seam · intermediate · ~16 min
  5. 05State in an immutable fleet - what survives replacement and what does notState · advanced · ~17 min
  6. 06The trade-offs, and the emergency change that has to existTrade-offs · advanced · ~16 min

Part LXXVI

Virtualisation and Linux

Linux as a VM: virtual hardware, virtio, VMware tools, cloud guest agents, CPU topology, memory ballooning, snapshots, time sync.

6 lessons
  1. 00Linux as a VM - virtual hardware, guest tools and the clockVirtual hardware · intermediate · ~14 min
  2. 01Guest agents - the out-of-band channel into a running VMGuest tooling · intermediate · ~16 min
  3. 02Converting a guest to virtio - the migration that makes a VM unbootableVirtual hardware · intermediate · ~12 min
  4. 03CPU topology in a guest, and what the guest cannot seeGuest visibility · advanced · ~17 min
  5. 04Timekeeping in a guest - clock sources, steps and the resumed VMGuest visibility · advanced · ~17 min
  6. 05Growing a guest online - disk, filesystem, memory and CPUGuest operations · intermediate · ~18 min

Part LXXVII

Linux in the Cloud

cloud-init, metadata services, ephemeral disks, persistent block storage, security groups, IAM concepts, instance lifecycle.

6 lessons
  1. 00cloud-init and the instance metadata serviceProvisioning · intermediate · ~14 min
  2. 01Ephemeral and persistent storage - the cloud storage modelEphemeral and persistent · intermediate · ~10 min
  3. 02Writing user-data - payload formats, module frequency, and the test loopProvisioning · intermediate · ~16 min
  4. 03Instance lifecycle - reboot, stop/start, terminate, replaceLifecycle · intermediate · ~16 min
  5. 04Security groups and the host firewall - two layers that fail differentlyNetwork boundary · intermediate · ~16 min
  6. 05Instance identity and cloud IAM from the Linux sideIdentity · intermediate · ~16 min

Part LXXVIII

Containers from the Linux Perspective

Namespaces, cgroups, OverlayFS, capabilities — the Linux primitives containers build on; cross-link to the Docker course.

6 lessons
  1. 01Container primitives - namespaces, cgroups, OverlayFS, capabilitiesPrimitives · advanced · ~12 min
  2. 02OverlayFS for containers - the layered filesystemOverlayFS · intermediate · ~10 min
  3. 03Network namespaces by hand - building container networking with ipNamespaces · advanced · ~18 min
  4. 04User namespaces and UID mapping - the kernel side of rootlessNamespaces · advanced · ~18 min
  5. 05What containers do not isolateBoundaries · advanced · ~17 min
  6. 06Attributing host symptoms to containers - triage without the runtime CLIOperations · advanced · ~18 min

Part LXXIX

Troubleshooting Methodology

Define symptom, determine impact, check recent changes, collect evidence, identify subsystem, form hypothesis, test safely, restore service, find root cause, prevent recurrence.

6 lessons
  1. 01The troubleshooting loop - ten steps and their stop-rulesThe loop · intermediate · ~12 minLab
  2. 02Evidence-based diagnosis - capture before you disturbThe loop · intermediate · ~12 minLab
  3. 03Preventing recurrence - closing the failure gap and the detection gapThe loop · intermediate · ~10 minLab
  4. 04Defining the symptom - turning a complaint into a measurementThe loop · intermediate · ~12 min
  5. 05What changed - working the highest-yield questionThe loop · intermediate · ~14 min
  6. 06Narrowing to one subsystem - bisect the path, then test one thingThe loop · intermediate · ~14 min

Part LXXX

Common Failure Scenarios

Service won’t start, system won’t boot, filesystem full, inode exhaustion, read-only FS, failed disk, degraded RAID, LVM full, DNS, route, packet loss, firewall, cert expiry, auth, sudo, OOM, CPU saturation, disk latency, broken fstab, kernel upgrade, cluster node loss, quorum loss.

6 lessons
  1. 01Failure: a service will not start, a host will not bootCommon failures · intermediate · ~12 min
  2. 02Failure: filesystem full, inodes exhausted, filesystem read-onlyCommon failures · intermediate · ~12 min
  3. 03Failure: DNS, routing, packet loss, firewall and certificate expiryCommon failures · intermediate · ~14 min
  4. 04Failure: OOM kills, CPU saturation and disk latencyCommon failures · intermediate · ~14 min
  5. 05Failure: cluster node loss, quorum loss and dependency outagesCommon failures · advanced · ~14 min
  6. 06Failure: nobody can log in, and sudo stopped workingCommon failures · intermediate · ~13 min

Part LXXXI

Incident Command

Impact assessment, severity, communication, roles, stabilisation, evidence preservation, recovery, escalation, post-incident review.

6 lessons
  1. 01Incident roles and communication - the team coordinationRoles and comms · intermediate · ~10 min
  2. 02Impact assessment - measuring what is broken for whomImpact assessment · intermediate · ~13 min
  3. 03Stabilisation - restoring service before you understand itStabilisation · intermediate · ~14 min
  4. 04Evidence preservation - what recovery destroysEvidence preservation · intermediate · ~14 min
  5. 05Escalation and handover - moving an incident between peopleEscalation · intermediate · ~13 min
  6. 06Recovery and stand-down - closing an incident properlyRecovery and stand-down · intermediate · ~13 min

Part LXXXII

Root Cause Analysis

Trigger vs contributing factor vs root cause, systemic causes, avoiding "human error" as the explanation.

6 lessons
  1. 01Avoiding human error as cause - the systemic approachHuman error · foundation · ~10 min
  2. 02Trigger, contributing factors, root cause - taking an incident apartDecomposition · intermediate · ~14 min
  3. 03Reconstructing the timeline - evidence over recollectionMethod · advanced · ~16 minLab
  4. 04Causal chains and stopping rules - not settling for the first plausible answerMethod · advanced · ~16 min
  5. 05Systemic causes - the conditions that made the incident possibleSystemic · expert · ~16 min
  6. 06The blameless postmortem - writing an RCA that produces changeOutput · advanced · ~15 min

Part LXXXIII

Operational Documentation

Architecture diagrams, inventories, build procedures, recovery procedures, dependency maps, maintenance procedures, escalation.

6 lessons
  1. 01Architecture and dependency maps - the documentation foundationArchitecture · foundation · ~10 min
  2. 02Build and recovery procedures - the operational referenceBuild and recovery · foundation · ~10 min
  3. 03Runbooks that work at 3amProcedures · intermediate · ~16 min
  4. 04Inventories, maintenance procedures and escalation pathsRecords · intermediate · ~15 min
  5. 05The incident record - writing it while it is happeningRecords · intermediate · ~16 min
  6. 06Documentation that stays true and can be foundMaintenance · intermediate · ~15 min

Part LXXXVI

Capstone: Production Linux Cluster

A 3-node production Linux cluster with HA load balancing, shared storage, monitoring, central logging, configuration management, central identity, DNS, NTP, backup, secrets, certificates — end-to-end operations.

0 lessons

No lessons published in this part yet. The full curriculum is planned in docs/courses/linux/curriculum.md on GitHub.

Part LXXXVII

Break It and Fix It

Major failure-injection section: symptoms first, evidence provided, root cause and remediation hidden behind reveal.

0 lessons