Linux · Curriculum
Curriculum
523 lessons across 85 parts. Lessons build on each other; later parts assume familiarity with earlier material.
Part I
Foundations
UNIX and Linux history, kernel vs userland, distributions, FHS, /proc /sys /dev /run /etc /var /usr /home /boot.
- 01Welcome to Linux for Production SysadminsCourse introduction · foundation · ~8 min
- 02History and ecosystem of UNIX and LinuxHistory · foundation · ~10 min
- 03Kernel vs userland — the layered modelArchitecture · foundation · ~8 min
- 04The Filesystem Hierarchy StandardFilesystem hierarchy · foundation · ~10 minLab
- 05/proc, /sys, /dev, /run — the kernel-exposed filesystemsKernel filesystems · foundation · ~8 min
- 06Linux distributions and support lifecyclesDistributions · foundation · ~6 min
Part II
Shell and Command-Line Operations
Bash fluency, quoting, expansion, redirection, pipelines, exit codes, core text tools, command composition.
- 01Shell anatomy — commands, options, arguments, exit codesShell fundamentals · foundation · ~8 min
- 02Quoting, expansion, and globbingShell expansion · foundation · ~12 min
- 03Environment variables, PATH, and shell initialisationEnvironment · foundation · ~10 min
- 04Redirection, pipes, and three-channel I/OPipelines and redirection · foundation · ~12 minLab
- 05Exit codes, conditionals, and testsControl flow · foundation · ~10 min
- 06Core text tools — grep, cut, sort, uniq, tr, wcText tools · foundation · ~12 minLab
- 07sed and awk in productionText tools · intermediate · ~12 min
- 08find, locate, and xargs — bulk operations on filesBulk operations · intermediate · ~10 minLab
- 09Shell history, aliases, and functionsProductivity · foundation · ~8 min
- 10jq and structured text — JSON, YAML, and CSV at the command lineStructured text · intermediate · ~8 min
Part III
Filesystems and Files
Inodes, hard/symbolic links, permissions, umask, ACLs, extended attributes, immutable attributes, timestamps.
- 01Linux file types and the seven file kindsFile types · foundation · ~8 min
- 02Inodes, dentries, and the VFS — how Linux tracks filesInodes · foundation · ~12 min
- 03Hard links and symbolic linksLinks · foundation · ~10 min
- 04Permissions, ownership, and umaskPermissions · foundation · ~12 min
- 05POSIX ACLs — fine-grained permissions beyond owner/group/otherPermissions · intermediate · ~12 min
- 06Extended attributes, immutable attributes, and capabilitiesAttributes · intermediate · ~10 min
Part IV
Users, Groups and Identity
/etc/passwd /etc/shadow /etc/group, UIDs, GIDs, supplementary groups, password policies, service accounts.
- 01/etc/passwd, /etc/shadow, /etc/groupIdentity databases · foundation · ~10 min
- 02UIDs, GIDs, and supplementary groupsIdentity · foundation · ~10 min
- 03Password policies and account lockingIdentity · intermediate · ~10 min
- 04Service accounts and system accountsIdentity · intermediate · ~10 min
- 05Name Service Switch (NSS) and getentIdentity · foundation · ~8 min
- 06Account lifecycle and offboarding - creating, changing and removing usersIdentity · intermediate · ~16 min
Part V
sudo and Privileged Access
Root vs sudo, sudoers, /etc/sudoers.d, least privilege, logging, privileged session management.
- 01Root vs sudo — the production privilege modelPrivileged access · foundation · ~8 min
- 02sudoers syntax — rules, aliases, and defaultssudoers · intermediate · ~12 min
- 03Drop-in files, package-installed rules, and #includedirsudoers · intermediate · ~8 min
- 04Least-privilege sudo designsudo design · intermediate · ~12 min
- 05Sudo logging, I/O capture, and SIEM integrationsudo logging · intermediate · ~8 min
- 06Shell escapes, NOEXEC, and sudoeditsudo design · intermediate · ~14 min
Part VI
Processes
PID/PPID, sessions, process groups, threads, states, zombies, signals, priorities, nice values, cgroup views.
- 01The Linux process model — PIDs, PPIDs, sessions, and groupsProcess model · foundation · ~10 min
- 02Process lifecycle — fork, exec, exit, and waitProcess lifecycle · intermediate · ~12 min
- 03Signals — SIGTERM, SIGKILL, SIGHUP, and the restSignals · intermediate · ~12 min
- 04Threads and cgroup viewsThreads · intermediate · ~8 min
- 05Nice, renice, and CPU scheduling prioritiesPriorities · foundation · ~10 min
- 06ps, pgrep, pkill, top, htop — the production toolkitObservation tools · foundation · ~10 min
Part VII
systemd and Service Management
PID 1, units, services, targets, timers, sockets, mounts, dependencies, drop-in overrides, resource controls.
- 01systemd architecture — PID 1 and the unit modelArchitecture · intermediate · ~12 min
- 02systemctl and unit managementUnit management · foundation · ~10 min
- 03Writing systemd unit filesUnit authoring · intermediate · ~14 min
- 04Targets, dependencies, and orderingTransactions · intermediate · ~10 min
- 05systemd timers — replacing cron with calendar and monotonic schedulesTimers · intermediate · ~10 min
- 06systemd sockets, paths, and activationActivation · intermediate · ~10 min
- 07Drop-in overrides — modifying units without rewritingDrop-ins · intermediate · ~8 min
- 08systemd-analyze, security review, and boot diagnosticsDiagnostics · intermediate · ~10 min
Part VIII
Logging and journald
Kernel logs, journald, syslog, rsyslog, logrotate, persistent journals, filtering, central logging.
- 01Linux logging architecture — kernel, journald, syslogLogging architecture · foundation · ~10 min
- 02journalctl — querying the systemd journalJournal queries · foundation · ~12 min
- 03rsyslog, syslog-ng, and traditional syslogrsyslog · intermediate · ~10 min
- 04logrotate — the rotation policylogrotate · intermediate · ~10 min
- 05Persistent journals and disk pressureJournal sizing · intermediate · ~8 min
- 06Centralised logging — rsyslog forwarding and journald remoteCentral logging · intermediate · ~12 min
Part IX
Boot Process
Firmware, UEFI, GRUB, initramfs, kernel parameters, emergency and rescue targets, recovery.
- 01Firmware, UEFI, and BIOSPre-boot · foundation · ~10 min
- 02GRUB — the bootloaderBootloader · intermediate · ~12 min
- 03Kernel, initramfs, and the first userlandKernel + initramfs · intermediate · ~12 min
- 04systemd targets and rescue modeTargets · intermediate · ~10 min
- 05Emergency and rescue mode — when normal boot failsEmergency · intermediate · ~12 min
- 06Boot failure diagnosis from the OOB consoleDiagnosis · intermediate · ~10 min
Part X
Kernel Management
Kernel versions, modules, dependencies, command line, sysctl, kernel logs, taint, live patching concepts.
- 01Kernel versions, releases, and lifecycleKernel versions · foundation · ~10 min
- 02Kernel modules — load, list, dependencies, blacklistModules · intermediate · ~12 min
- 03sysctl — runtime kernel tuningsysctl · intermediate · ~10 min
- 04Kernel logs, taint flags, and oops decodingKernel logs · intermediate · ~10 min
- 05Live kernel patching — kpatch, kgraft, and livepatchLive patching · advanced · ~8 min
- 06Kernel command line parameters - the settings you can only make at bootKernel command line · intermediate · ~14 min
Part XI
Package Management
apt, dpkg, dnf, rpm, repositories, signing, dependency resolution, version pinning, rollback considerations.
- 01Package management overview — apt vs dnfPackage management · foundation · ~15 min
- 02apt and dpkg — Debian-family package managementapt and dpkg · foundation · ~12 min
- 03dnf and rpm — RHEL-family package managementdnf and rpm · foundation · ~12 min
- 04Repositories, signing, and trusted sourcesRepositories · intermediate · ~10 min
- 05Package pinning, holds, and version selectionVersion pinning · intermediate · ~10 min
- 06Security updates and CVE responseSecurity updates · intermediate · ~10 min
- 07Package cache, cleanup, and rollbackCache and rollback · foundation · ~8 min
Part XII
Repository Security and Supply Chain
GPG keys, package provenance, third-party repositories, malicious packages, dependency compromise.
- 01Repository trust and GPG signingTrust · foundation · ~10 min
- 02Third-party repository riskThird-party repos · intermediate · ~10 min
- 03Package provenance and SBOMProvenance · intermediate · ~10 min
- 04Mitigating curl | bash and other supply-chain footgunsSupply chain · intermediate · ~8 min
- 05Malicious packages and maintainer compromiseSupply chain · advanced · ~16 min
- 06Language package managers on production hostsSupply chain · advanced · ~16 min
Part XIII
Disks and Block Devices
Block devices, naming, sectors, partitions, GPT, signatures, UUIDs, labels, persistent naming.
- 01Block devices, naming, and disk identificationBlock devices · foundation · ~10 min
- 02Partitions, GPT, and MBRPartitioning · foundation · ~10 min
- 03Filesystem signatures and wipefsFilesystem signatures · foundation · ~8 min
- 04Persistent naming - UUIDs, labels, and stable pathsPersistent naming · foundation · ~8 min
- 05Sector sizes and partition alignment - 512n, 512e, and 4KnSectors and alignment · intermediate · ~15 min
- 06Rescanning and growing block devices - when the kernel still sees the old sizeRescan and resize · advanced · ~15 min
Part XIV
Filesystems
ext4, XFS, Btrfs characteristics, journaling, mount options, online growth, inode exhaustion, fsck/xfs_repair.
- 01Filesystem concepts - journaling, allocation, and durabilityFilesystem concepts · foundation · ~10 min
- 02ext4 deep dive - features, tuning, and operationsext4 · intermediate · ~12 min
- 03XFS deep dive - features, tuning, and operationsxfs · intermediate · ~12 min
- 04btrfs where it fits - snapshots, subvolumes, and operationsbtrfs · intermediate · ~10 min
- 05Inode exhaustion - the filesystem that says noOperations · intermediate · ~8 min
- 06Filesystem check, repair, and recovery with fsck and xfs_repairOperations · intermediate · ~10 min
Part XV
/etc/fstab and Mount Management
UUID, LABEL, mount options, systemd interaction, network filesystems, recovery from broken fstab.
- 01fstab syntax - fields, options, and the dump/pass columnsfstab syntax · foundation · ~10 min
- 02Mount options and systemd integrationMount options · intermediate · ~10 min
- 03Recovering from a broken fstab without a live USBBoot recovery · intermediate · ~10 min
- 04Network filesystems in fstab - _netdev, automount, and the boot you cannot watchNetwork filesystems · intermediate · ~15 min
- 05fstab entries that are not disks - bind mounts, tmpfs, and swapNon-device entries · intermediate · ~15 min
- 06Mount units by hand, and verifying a mount change before you rebootMount units · advanced · ~15 min
Part XVI
LVM
PV/VG/LV, extension, snapshots, risks of shrinking, recovery, day-2 operations.
- 01LVM concepts - PV, VG, LV, and the storage stackLVM concepts · foundation · ~10 min
- 02LVM creation, extension, and snapshotsLVM operations · intermediate · ~12 min
- 03LVM snapshots and rollback - the local safety netSnapshots · intermediate · ~8 min
- 04LVM monitoring and capacity planningMonitoring · intermediate · ~8 min
- 05LVM troubleshooting - common failures and how to recoverTroubleshooting · intermediate · ~8 min
- 06Shrinking an LV - the highest-risk LVM operationLVM operations · advanced · ~10 min
- 07LVM thin provisioning and thin-pool monitoringMonitoring · advanced · ~14 min
Part XVII
Software RAID
mdadm, RAID1/5/6/10, degraded arrays, rebuilds, failure detection, monitoring.
- 01mdadm RAID concepts - levels, layouts, and failure modesRAID concepts · foundation · ~10 min
- 02RAID1 mirroring - simple and reliableRAID levels · foundation · ~8 min
- 03RAID5 and RAID6 - parity for capacity and fault toleranceRAID levels · intermediate · ~10 min
- 04RAID10 - mirrored stripes for performance and redundancyRAID levels · intermediate · ~8 min
- 05Degraded arrays and rebuilds - replacing a failed diskRAID operations · intermediate · ~8 min
- 06md RAID monitoring and failure detectionRAID operations · intermediate · ~14 min
Part XVIII
Enterprise Storage
SAN, NAS, iSCSI, Fibre Channel, NFS, multipath, device-mapper, shared storage failure paths.
- 01SAN vs NAS vs object storage - which one to chooseStorage taxonomy · foundation · ~8 min
- 02iSCSI - block storage over IPiSCSI · intermediate · ~10 min
- 03Fibre Channel - high-speed block storageFC · intermediate · ~8 min
- 04NFS client - mounting and operating network filesystemsNFS · intermediate · ~10 min
- 05Multipath and device-mapper - redundant paths to SAN LUNsMultipath · intermediate · ~8 min
- 06Shared storage failure paths - reading a host whose storage has gone awayFailure paths · advanced · ~18 min
Part XIX
Networking Foundations
Ethernet, MAC, ARP, IPv4, IPv6, subnetting, routing, TCP, UDP, ICMP, DNS, DHCP concepts.
- 01OSI and TCP/IP models - what each layer is forModels · foundation · ~10 min
- 02Ethernet and MAC addressing - what is on the wireEthernet · foundation · ~10 min
- 03IPv4 addressing, subnetting, and CIDRIPv4 · foundation · ~14 min
- 04IPv6 addressing and subnettingIPv6 · foundation · ~12 min
- 05ARP and IPv6 neighbour discoveryARP/ND · foundation · ~10 min
- 06TCP and UDP - transport-layer essentialsTCP/UDP · foundation · ~14 min
- 07ICMP and path MTU - control messages and fragmentationICMP/MTU · foundation · ~10 min
- 08DNS and DHCP conceptsDNS/DHCP · foundation · ~12 min
Part XX
Linux Network Configuration
Interfaces, addresses, routes, gateways, DNS, Netplan, NetworkManager, systemd-networkd, persistent config.
- 01The ip command - the sysadmin network toolkitip command · foundation · ~12 min
- 02Linux interface configuration - addresses, MTU, offloadsInterface config · foundation · ~12 min
- 03Routing and default gateways on LinuxRouting · foundation · ~12 min
- 04Netplan on Ubuntu - declarative network configurationNetplan · foundation · ~10 min
- 05NetworkManager - flexible configuration for desktops and laptopsNetworkManager · foundation · ~10 min
- 06systemd-networkd - predictable configuration for serverssystemd-networkd · foundation · ~10 min
- 07Distribution comparison - network configuration across distrosComparison · foundation · ~8 min
Part XXI
Advanced Linux Networking
VLANs, bridges, bonding, teaming, MTU, jumbo frames, multiple routing tables, policy routing, IPv6, LACP.
- 01VLANs and 802.1Q - segmenting a network with tagsVLAN · intermediate · ~12 min
- 02Linux bridges and bridge VLANsBridge · intermediate · ~12 min
- 03Bonding and LACP - aggregating links for redundancy and throughputBond · intermediate · ~14 min
- 04NIC teaming - the userspace alternative to bondingTeaming · intermediate · ~8 min
- 05MTU and jumbo frames - when 1500 is not enoughMTU · intermediate · ~10 min
- 06Multiple routing tables on LinuxRouting tables · advanced · ~12 min
- 07Policy routing - choosing routes by source, mark, or interfacePolicy routing · advanced · ~12 min
- 08IPv6 advanced - privacy, SLAAC details, and operational gotchasIPv6 advanced · advanced · ~10 min
Part XXII
Network Troubleshooting
OSI/TCP-IP layer methodology, ip, ss, mtr, dig, tcpdump, ethtool, arping, nc, openssl s_client, evidence-based diagnosis.
- 01Network troubleshooting methodology - the systematic approachMethodology · foundation · ~14 min
- 02ss and socket state - inspecting TCP and UDP from the kernelss · foundation · ~12 min
- 03ping, traceroute, and mtr - path diagnosisPath diagnosis · foundation · ~12 min
- 04dig and nslookup - DNS diagnosis in detailDNS · foundation · ~12 min
- 05tcpdump and packet capture - seeing the bytestcpdump · intermediate · ~14 min
- 06ethtool and offloads - link-layer diagnosticsethtool · intermediate · ~12 min
- 07arping and ARP inspection - forcing layer-2 discoveryARP · intermediate · ~8 min
- 08nc and openssl s_client - probing TCP and TLSProbing · intermediate · ~12 min
- 09curl as a diagnostic tool - HTTP, headers, and timingcurl · foundation · ~10 min
Part XXIII
DNS
Resolver architecture, /etc/resolv.conf, systemd-resolved, recursive resolution, authoritative DNS, TTL, caching, cluster failure modes.
- 01DNS resolver architecture - how Linux finds namesArchitecture · foundation · ~12 min
- 02/etc/resolv.conf - the resolver configuration fileresolv.conf · foundation · ~10 min
- 03systemd-resolved - the modern stub resolversystemd-resolved · intermediate · ~15 min
- 04Recursive resolution - how a name becomes an IPResolution · foundation · ~10 min
- 05Authoritative DNS, TTLs, and caching behaviourAuthoritative · intermediate · ~12 min
- 06Cluster DNS failure modes - what breaks when DNS is downCluster DNS · advanced · ~10 min
Part XXIV
Time Synchronisation
UTC, monotonic time, NTP, chrony, systemd-timesyncd, operational impact on auth, TLS, logs, distributed systems.
- 01Time concepts - UTC, monotonic, and the wall clockConcepts · foundation · ~10 min
- 02chrony architecture - the modern NTP client and serverchrony · intermediate · ~12 min
- 03chrony client configuration - tuning for accuracy and resiliencechrony config · intermediate · ~12 min
- 04systemd-timesyncd - the lightweight NTP clientsystemd-timesyncd · intermediate · ~8 min
- 05Time skew operational impact - what breaks when clocks driftOperational impact · intermediate · ~10 min
- 06NTP fleet design - hierarchy, falsetickers and leap secondsFleet design · advanced · ~15 min
Part XXV
Firewalls
nftables as the modern foundation, plus iptables/firewalld/ufw differences, chains, hooks, stateful filtering, connection tracking, NAT.
- 01Firewall concepts - what a Linux firewall actually doesConcepts · foundation · ~10 min
- 02nftables architecture - tables, chains, and rulesnftables · intermediate · ~14 min
- 03nftables chains, tables, and hooks in depthnftables chains · intermediate · ~12 min
- 04nftables stateful filtering and conntracknftables stateful · intermediate · ~12 min
- 05nftables NAT - source NAT, destination NAT, and masqueradenftables NAT · intermediate · ~12 min
- 06iptables legacy and migration to nftablesiptables · intermediate · ~10 min
- 07firewalld - zone-based firewall managementfirewalld · intermediate · ~10 min
- 08ufw and distribution firewall wrappersufw · intermediate · ~8 min
- 09Validating actual firewall behaviourValidation · intermediate · ~10 min
Part XXVI
SSH
Client/server, keys, host keys, agent, forwarding, ProxyJump, configuration, certificates, MFA concepts, bastion hosts.
- 01SSH architecture - how Secure Shell actually worksArchitecture · foundation · ~10 min
- 02SSH keys and known_hosts - the key-based authentication modelKeys · foundation · ~12 min
- 03ssh-agent and forwarding - keeping keys safeAgent · intermediate · ~10 min
- 04sshd configuration - hardening the serversshd config · intermediate · ~14 min
- 05SSH certificates and bastionsCertificates · intermediate · ~12 min
- 06ProxyJump and jump hosts - secure access patternsProxyJump · intermediate · ~10 min
- 07SSH hardening - production checklistHardening · intermediate · ~12 min
- 08SSH MFA and 2FA conceptsMFA · intermediate · ~10 min
Part XXVII
Authentication and Enterprise Identity
PAM, NSS, LDAP, Active Directory, Kerberos, SSSD, central identity failure modes.
- 01PAM architecture - pluggable authentication on LinuxPAM · intermediate · ~12 min
- 02NSS - Name Service Switch for users, groups, and hostsNSS · intermediate · ~10 min
- 03LDAP and Active Directory - directory services fundamentalsLDAP · intermediate · ~12 min
- 04Kerberos concepts - tickets, realms, and principalsKerberos · advanced · ~12 min
- 05SSSD architecture - the system security services daemonSSSD · advanced · ~12 min
- 06SSSD and Active Directory integrationSSSD AD · advanced · ~14 min
- 07Central identity failure modes - what breaks when LDAP or AD is downIdentity failure · advanced · ~10 min
- 08PAM and NSS troubleshooting - fixing auth that breaksPAM NSS debug · intermediate · ~10 min
Part XXVIII
SELinux and AppArmor
Mandatory access control, contexts, policies, enforcement, audit logs, denial investigation, profile creation.
- 01MAC vs DAC - mandatory versus discretionary access controlConcepts · foundation · ~10 min
- 02SELinux architecture - contexts, policy, and enforcementSELinux · advanced · ~12 min
- 03SELinux contexts and labels - inspecting and fixingSELinux contexts · advanced · ~10 min
- 04SELinux policy and modes - enforcing, permissive, and custom policiesSELinux policy · advanced · ~12 min
- 05AppArmor concepts - path-based MAC for Debian-family systemsAppArmor · intermediate · ~10 min
- 06MAC denial investigation - diagnosing SELinux and AppArmor denialsMAC debug · intermediate · ~10 min
Part XXIX
Linux Security Hardening
Least privilege, SSH, sudo, mount options, kernel parameters, firewall, MAC, audit, packages, CIS-style controls.
- 01Linux hardening principles - defence in depthPrinciples · foundation · ~10 min
- 02SSH and sudo hardening - the privileged access surfaceSSH sudo · intermediate · ~12 min
- 03Kernel hardening with sysctl - the runtime network and security parameterssysctl · intermediate · ~10 min
- 04Mount options for security - nodev, nosuid, noexecMount options · foundation · ~8 min
- 05Service minimisation - reducing the attack surfaceService minimisation · intermediate · ~10 min
- 06Package update discipline - patching as a security controlUpdates · foundation · ~10 min
- 07CIS benchmarks conceptually - mapping controls to the frameworkCIS · intermediate · ~10 min
Part XXX
Linux Capabilities and Privilege
Traditional root model, Linux capabilities, file capabilities, process capabilities, capsh, getcap, setcap.
- 01Traditional root model - why capabilities existRoot model · foundation · ~10 min
- 02Linux capabilities overview - the full list and categoriesCapabilities overview · intermediate · ~12 min
- 03File capabilities - replacing setuid with fine-grained privilegesFile capabilities · intermediate · ~12 min
- 04capsh and capability debugging - inspecting and dropping capabilitiesCapability debug · intermediate · ~10 min
- 05Process capability sets and the exec transitionProcess capabilities · advanced · ~16 min
- 06Root-equivalent capabilities - when dropping to one changes nothingProcess capabilities · advanced · ~15 min
Part XXXI
Audit and Security Logging
auditd, authentication logs, sudo logs, kernel security events, ausearch, aureport, auditctl, central security logging.
- 01auditd architecture - the Linux audit frameworkauditd · intermediate · ~12 min
- 02audit rules and watch points - what to recordAudit rules · intermediate · ~12 min
- 03ausearch and aureport - querying the audit logAusearch aureport · foundation · ~10 min
- 04Authentication and sudo logs - what gets recorded and how to read itAuth logs · foundation · ~10 min
- 05Central security logging - shipping logs to a SIEMCentral logging · intermediate · ~10 min
- 06Audit record anatomy and attribution - answering whoAudit analysis · advanced · ~16 min
Part XXXII
Vulnerability and Patch Management
Detection, CVEs, severity vs context, exploitability, patch priority, maintenance windows, reboot requirements.
- 01Vulnerability concepts - CVEs, severity, and exploitabilityConcepts · foundation · ~10 min
- 02CVE severity vs context - prioritising for the environmentSeverity vs context · intermediate · ~10 min
- 03Scanning and detecting vulnerabilitiesScanning · intermediate · ~12 min
- 04Patch priority decisions - timing and riskPatch priority · intermediate · ~10 min
- 05Emergency patching - reacting to critical CVEs under pressureEmergency · advanced · ~10 min
- 06Reboot and restart requirements - finishing the patchPatch execution · intermediate · ~15 min
Part XXXIII
Fleet Patch Management
Dev/test/canary/production waves, staged deployments, health validation, rollback, automated patching trade-offs.
- 01Fleet patch strategy - waves, canaries, and rollbackStrategy · intermediate · ~10 min
- 02Canary and wave deployments - phased rollout with health checksCanary waves · intermediate · ~10 min
- 03Rollback strategies - safe deployment reversalsRollback · intermediate · ~10 min
- 04Pre-patching and post-patching validationValidation · intermediate · ~10 min
- 05Automated patching in a fleet - what to automate and what never toAutomation · advanced · ~16 min
- 06Repository snapshots - making a wave reproducibleAutomation · advanced · ~17 min
Part XXXIV
Configuration Management
Desired state, idempotency, drift, inventory, automation, Ansible-flavoured examples.
- 01Why manual does not scale - the case for configuration managementWhy CM · foundation · ~10 min
- 02Desired state and idempotency - the foundations of CMDesired state · intermediate · ~10 min
- 03Ansible for sysadmins - the modern CM toolAnsible · intermediate · ~12 min
- 04Inventory and facts - knowing your fleetInventory · intermediate · ~10 min
- 05Config drift detection - keeping reality aligned with declared stateDrift detection · intermediate · ~10 min
- 06Safe CM rollout - blast radius, check mode, and the control nodeSafe rollout · advanced · ~15 min
Part XXXV
Shell Scripting for Sysadmins
Variables, conditions, loops, functions, traps, exit codes, error handling, set -euo pipefail, ShellCheck.
- 01Script structure and conventions - the production shell scriptStructure · intermediate · ~10 min
- 02Variables and quoting - the foundation of safe shell scriptsVariables · intermediate · ~10 min
- 03Conditions and loops - controlling script flowConditions loops · foundation · ~10 min
- 04Functions and arguments - reusable, testable script componentsFunctions · intermediate · ~10 min
- 05Error handling - set -euo pipefail and trapError handling · intermediate · ~12 min
- 06Traps and temporary files - clean exit on signal or errorTraps tmpfiles · intermediate · ~10 min
- 07ShellCheck and static analysis - catching bugs before runtimeShellCheck · foundation · ~10 min
Part XXXVI
Scheduled Operations
cron, anacron, systemd timers, locking, duplicate prevention, logging, failure detection.
- 01cron syntax and anacron - the classic schedulercron · foundation · ~10 min
- 02systemd timers replacing cron - the modern schedulersystemd timers · intermediate · ~12 min
- 03Locking and concurrency - preventing duplicate jobsLocking · intermediate · ~10 min
- 04Job logging and failure detection - knowing what happenedLogging · intermediate · ~10 min
- 05Timeouts, retries and jitter - scheduled jobs that fail safelyJob safety · intermediate · ~18 min
- 06Auditing the schedule - what actually runs on this hostSchedule inventory · intermediate · ~17 min
Part XXXVII
Resource Management
CPU, memory, swap, processes, file descriptors, ulimits, cgroups, systemd resource controls, OOM behaviour.
- 01cgroups v2 architecture - the Linux resource control subsystemcgroups · advanced · ~12 min
- 02ulimits and limits.conf - per-user resource limitsulimits · foundation · ~10 min
- 03File descriptors and /proc/fd - understanding open filesFile descriptors · intermediate · ~14 min
- 04systemd resource controls - applying cgroups via unitssystemd resource controls · intermediate · ~12 min
- 05OOM killer behaviour - what happens when memory runs outOOM killer · advanced · ~10 min
- 06cgroup accounting - proving which limit was actually hitResource accounting · advanced · ~20 min
Part XXXVIII
Linux Performance Fundamentals
USE methodology, top, htop, vmstat, mpstat, pidstat, iostat, sar, free, slabtop, evidence-based investigation.
- 01USE methodology - the framework for performance investigationUSE method · foundation · ~10 min
- 02top, htop, and vmstat - the basic performance toolstop htop vmstat · foundation · ~10 min
- 03mpstat and pidstat - per-CPU and per-process statisticsmpstat pidstat · intermediate · ~10 min
- 04iostat and sar - storage and historical performanceiostat sar · intermediate · ~10 min
- 05free, slabtop, and /proc/meminfo - memory diagnosticsfree slabtop meminfo · intermediate · ~10 min
- 06Measure, change one thing, measure again - the tuning disciplineMethod · intermediate · ~19 min
Part XXXIX
CPU Performance
Utilisation, load average, run queue, context switching, interrupts, softirqs, steal time, CPU affinity, NUMA.
- 01CPU utilisation and load average - measuring the workloadCPU utilisation · foundation · ~10 min
- 02Run queue and context switching - measuring contentionRun queue · intermediate · ~10 min
- 03Interrupts, softirqs, and steal time - hardware and VM signalsInterrupts steal · intermediate · ~10 min
- 04CPU affinity and NUMA - controlling where processes runCPU affinity · intermediate · ~10 min
- 05CPU frequency, governors and thermal throttling - the clock is not constantFrequency · advanced · ~19 min
- 06perf - finding the code that is burning the CPUProfiling · advanced · ~21 min
Part XL
Memory Performance
Virtual memory, pages, page cache, anonymous memory, buffers, swap, memory pressure, OOM killer, slab.
- 01Virtual memory and pages - how Linux manages memoryVirtual memory · intermediate · ~10 min
- 02Page cache and anonymous memory - the two memory kindsPage cache anon · intermediate · ~10 min
- 03Swap and memory pressure - when the system runs out of RAMSwap · intermediate · ~10 min
- 04OOM killer decisions - who gets killed and whyOOM killer decisions · advanced · ~10 min
- 05RSS, PSS and shared pages - what a process really costsProcess footprint · advanced · ~19 min
- 06Kernel memory, huge pages and the tuning knobs that backfireKernel memory · advanced · ~21 min
Part XLI
Storage Performance
Latency, throughput, IOPS, queue depth, utilisation, filesystem cache, fio, iostat, iotop, safe benchmarking.
- 01I/O latency, throughput, and IOPS - the three storage metricsI/O metrics · foundation · ~10 min
- 02Queue depth and utilisation - storage saturation signalsQueue depth · intermediate · ~16 min
- 03iotop and pidstat - per-process I/O statisticsiotop pidstat · intermediate · ~10 min
- 04fio and safe benchmarking - measuring storage baselinefio · intermediate · ~12 min
- 05The I/O stack - which layer is adding the latencyI/O stack · advanced · ~21 min
- 06Writeback and filesystem tuning - dirty pages, mount options and fstrimFilesystem tuning · advanced · ~20 min
Part XLII
Network Performance
Bandwidth, latency, packet loss, retransmissions, socket queues, connection states, iperf3, ethtool, sar.
- 01Bandwidth and latency - the two network metricsBandwidth latency · foundation · ~10 min
- 02Packet loss and retransmits - when the network dropsPacket loss · intermediate · ~10 min
- 03Socket queues and connection states - TCP tuningSocket queues · intermediate · ~10 min
- 04iperf3 and throughput testing - measuring bandwidthiperf3 · foundation · ~10 min
- 05The bandwidth-delay product - why a fast link runs slowLong fat networks · advanced · ~20 min
- 06sar -n - network history and the counters nobody collectedHistorical data · advanced · ~18 min
Part XLIII
eBPF and Advanced Observability
eBPF, tracepoints, kprobes, uprobes, BPF maps, bpftrace, BCC, sysadmin use cases.
- 01eBPF overview - in-kernel observability for sysadminseBPF · advanced · ~12 min
- 02Tracepoints, kprobes, and uprobes - eBPF attachment pointsTracepoints kprobes · advanced · ~10 min
- 03bpftrace and bcc - sysadmin use cases and examplesbpftrace bcc · advanced · ~10 min
- 04BPF maps and state - communicating with userspaceBPF maps · advanced · ~10 min
- 05Running eBPF safely - kernel requirements, permissions and lockdownRequirements · expert · ~21 min
- 06The cost of tracing - eBPF overhead, dropped events and production safetyProduction safety · expert · ~21 min
Part XLIV
Central Monitoring
Prometheus node_exporter, Grafana, alerting on CPU, memory, FS, inodes, disk latency, network, services, time sync, hardware.
- 01Monitoring design - what to monitor, how to alert, and whyMonitoring design · foundation · ~10 min
- 02Prometheus and node_exporter - the standard Linux monitoring stackPrometheus · intermediate · ~10 min
- 03Grafana and alerting - dashboards and notificationGrafana alerting · intermediate · ~10 min
- 04What to monitor on Linux - the production checklistWhat to monitor · foundation · ~10 min
- 05Hardware monitoring sensors - temperature, voltage, fanHardware sensors · foundation · ~10 min
- 06Operating the monitoring stack - retention, cardinality and who watches the watcherMonitoring design · advanced · ~15 min
Part XLV
Central Logging
journald, rsyslog, syslog, Fluent Bit, Vector, Loki, Elasticsearch, retention, filtering, cardinality, capacity.
- 01Central logging design - the architecture and decisionsLogging design · foundation · ~10 min
- 02rsyslog forwarding - traditional syslog shippingrsyslog · intermediate · ~10 min
- 03Fluent Bit and Vector - modern log shippersFluent Bit Vector · intermediate · ~10 min
- 04Loki and Elasticsearch - the modern log storeLoki Elasticsearch · intermediate · ~10 min
- 05Log cardinality and capacity - sizing and tuningCardinality capacity · advanced · ~10 min
- 06Where log messages go missing - and how to prove they did notCentral logging · advanced · ~15 min
Part XLVI
OpenTelemetry
Metrics, logs, traces, collectors, agents/gateways, Linux infrastructure telemetry.
- 01OpenTelemetry three pillars - metrics, logs, tracesOTel three pillars · foundation · ~10 min
- 02OTel collectors and agents - the deployment modelOTel collectors · intermediate · ~10 min
- 03OTel host receiver and process metricsOTel host receiver · intermediate · ~10 min
- 04OTel into Grafana and Tempo - the visual stackOTel visualisation · intermediate · ~10 min
- 05OTel logs on a Linux host - the journald and filelog receiversOTel collectors · advanced · ~14 min
- 06The cost of telemetry - cardinality, sampling and collector limitsOTel collectors · advanced · ~15 min
Part XLVII
Backup Strategy
Filesystem, application, database, configuration, snapshot, consistency, encryption, retention, immutable copies, offsite, 3-2-1.
- 01Backup concepts - what to back up and howConcepts · foundation · ~10 min
- 02Application-consistent vs crash-consistent backupsConsistency · intermediate · ~10 min
- 033-2-1 and modern backup strategies3-2-1 · foundation · ~10 min
- 04Immutable and offline copies - the last line of defenceImmutable offline · intermediate · ~10 min
- 05Encryption and key management - protecting backups at restEncryption keys · advanced · ~10 min
- 06Backing up what is not a file - partition tables, LVM metadata, LUKS headers and package stateStructural metadata · advanced · ~18 min
Part XLVIII
Backup Tools
rsync, tar, Borg, Restic, filesystem snapshots, enterprise backup integration, selection criteria.
- 01rsync and tar - the classic backup toolsrsync tar · foundation · ~10 min
- 02BorgBackup and Restic - the modern deduplicated backup toolsBorg Restic · intermediate · ~12 min
- 03Filesystem snapshots as backup - ZFS and btrfsFS snapshots · intermediate · ~10 min
- 04LVM snapshots for backups - the classic Linux approachLVM snapshots · intermediate · ~10 min
- 05Enterprise backup integration - Veeam, NetBackup, and similarEnterprise · intermediate · ~10 min
- 06Backup tool selection - choosing the right tool for the workloadTool selection · intermediate · ~10 min
Part XLIX
Restore
Restoration over backup success: files, ownership, ACLs, services, configuration, full restore exercises.
- 01Restore matters more than backup - the discipline of restoreRestore matters · foundation · ~10 min
- 02Restore testing discipline - the regular validation routineRestore testing · intermediate · ~10 min
- 03Restore files and ownership - the practical detailsFiles ownership · intermediate · ~10 min
- 04Restore services and configuration - the full-system recoveryServices config · advanced · ~10 min
- 05Partial and point-in-time restores - one file, one directory, one momentPartial restore · advanced · ~18 min
- 06Restore drills - turning an asserted RTO into a measured oneDrills and measurement · advanced · ~18 min
Part L
Disaster Recovery
RPO, RTO, disaster scenarios, rebuild vs restore, bare-metal recovery, DNS, certificates, identity, full cluster loss.
- 01RPO and RTO - modelling recovery objectivesRPO RTO · foundation · ~10 min
- 02Disaster scenarios - what to plan forDisaster scenarios · foundation · ~10 min
- 03Bare-metal recovery - restoring to a fresh hostBare metal · advanced · ~10 min
- 04Dependencies: DNS, certificates, identity - the cascade risksDependencies · advanced · ~10 min
- 05Full cluster loss recovery - the worst-case scenarioFull cluster loss · advanced · ~10 min
- 06Declaring a disaster, failing over, and failing backDeclaration and failback · expert · ~20 min
Part LI
Linux Fleet Architecture
Management plane, configuration management, monitoring, logging, identity, automation, secrets, patching across 50-5000 nodes.
- 01Fleet management plane - how a Linux fleet is operatedManagement plane · foundation · ~10 min
- 02Naming and inventory - knowing your fleetNaming inventory · foundation · ~10 min
- 03Environment separation - dev, staging, and productionEnvironment separation · foundation · ~10 min
- 04Central identity and secrets - the foundation of fleet securityCentral identity secrets · advanced · ~10 min
- 05Host enrolment and decommissioning - joining and leaving the fleetFleet lifecycle · advanced · ~16 min
- 06Scaling the management plane - what breaks between 50 and 5000 nodesFleet lifecycle · expert · ~17 min
Part LII
High Availability Fundamentals
Availability, redundancy, fault tolerance, failure domains, active/active, active/passive, N+1, N+2, host vs service availability.
- 01Availability and reliability - the metrics that matterAvailability · foundation · ~10 min
- 02Redundancy and fault tolerance - building for failureRedundancy · foundation · ~10 min
- 03Failure domains - what fails togetherFailure domains · intermediate · ~10 min
- 04N+1 capacity - the math of headroomN+1 · intermediate · ~10 min
- 05Host availability vs service availability - measuring the right thingHost vs service · intermediate · ~12 min
- 06Redundancy levels - N+1, N+2, 2N and 2N+1Redundancy levels · intermediate · ~12 min
Part LIII
Quorum and Split Brain
Quorum, majority, membership, partitions, split brain, witness/quorum device, failure detection.
- 01Quorum concepts - the math of agreementQuorum · advanced · ~10 min
- 02Membership and partitions - the failure mode that breaks quorumMembership · advanced · ~10 min
- 03Split brain explained - the cluster failure modeSplit brain · advanced · ~10 min
- 04Witness and quorum devices - breaking geographic tiesWitness · advanced · ~10 min
- 05Failure detection - how a cluster decides a node has failedFailure detection · advanced · ~13 min
- 06Two-node clusters and tie-breaking - the honest versionTwo-node clusters · advanced · ~14 min
Part LIV
Fencing and STONITH
Why fencing exists, split-brain data corruption, fencing devices, STONITH, hardware-independent principles.
- 01Why fencing exists - the cluster disciplineFencing · advanced · ~10 min
- 02Fencing devices and agents - the practical implementationFencing devices · advanced · ~10 min
- 03STONITH and data integrity - the shoot-the-other-node-in-the-head patternSTONITH · advanced · ~10 min
- 04Unresponsive is not dead - the evidence problem fencing solvesEvidence · advanced · ~14 min
- 05SBD and watchdog fencing - when the node fences itselfSBD · advanced · ~14 min
- 06Fencing design principles - evaluating any platformDesign principles · advanced · ~13 min
Part LV
Pacemaker and Corosync
Corosync membership, Pacemaker, resources, resource agents, constraints, failover, fencing integration.
- 01Corosync architecture - the cluster membership layerCorosync · advanced · ~10 min
- 02Pacemaker architecture - the resource managerPacemaker · advanced · ~10 min
- 03Pacemaker resources and resource agentsPacemaker resources · advanced · ~10 min
- 04Pacemaker constraints: location, colocation, orderPacemaker constraints · advanced · ~10 min
- 05Pacemaker fencing integration - tying STONITH to resourcesPacemaker fencing · advanced · ~10 min
- 06Pacemaker troubleshooting - the systematic approachTroubleshooting · advanced · ~10 min
Part LVI
Keepalived and VRRP
Virtual IPs, VRRP, health checks, master/backup, failover, limitations.
- 01VRRP concepts - virtual router redundancy protocolVRRP · foundation · ~10 min
- 02Keepalived configuration - the Linux VRRP implementationKeepalived config · intermediate · ~10 min
- 03Keepalived health checks - service-aware failoverHealth checks · intermediate · ~10 min
- 04Keepalived limitations - when VRRP is not enoughLimitations · intermediate · ~10 min
- 05Virtual IP mechanics - what actually moves during a failoverVIP mechanics · intermediate · ~14 min
- 06Sync groups and multiple VIPs - keeping a service togetherSync groups · intermediate · ~13 min
Part LVII
Linux Load Balancing
HAProxy, nginx, IPVS, Layer 4 vs Layer 7, health checking, algorithms, session persistence, connection draining.
- 01Load balancing concepts - distributing traffic across backendsConcepts · foundation · ~10 min
- 02HAProxy configuration - the production-grade layer 7 LBHAProxy · intermediate · ~12 min
- 03nginx as a load balancer - the versatile optionnginx LB · intermediate · ~10 min
- 04IPVS and the kernel load balancer - the layer 4 workhorseIPVS · advanced · ~10 min
- 05Health checks and persistence - the safety net of load balancingHealth checks · intermediate · ~10 min
- 06Connection draining - taking a backend out without dropping requestsLoad balancing · advanced · ~14 min
Part LVIII
Clustered Service Architecture
Stateless, shared state, replicated state, external state, application architecture constraints, identifying the right pattern.
- 01Stateless services - the simplest cluster architectureStateless · foundation · ~10 min
- 02Shared state and replicated state - when stateless is not enoughShared state · advanced · ~10 min
- 03External state and databases - the stateful backendExternal state · advanced · ~10 min
- 04Application architecture constraints - the limits of clusteringConstraints · advanced · ~10 min
- 05Identifying the right cluster pattern - a decision procedureChoosing a pattern · advanced · ~13 min
- 06Failover as the client sees it - the outage the cluster does not measureClient-visible recovery · advanced · ~13 min
Part LX
Distributed Storage Concepts
Ceph, Gluster, object storage, replicated block, replication, quorum, failure domains, consistency, recovery.
- 01Ceph architecture intro - the modern distributed storageCeph · advanced · ~10 min
- 02Gluster concepts intro - the other distributed filesystemGluster · intermediate · ~10 min
- 03Object and block storage tradeoffs - the right tool for the dataObject vs block · intermediate · ~10 min
- 04Replication, quorum, and consistency - the distributed storage disciplineReplication · advanced · ~10 min
- 05Failure domains in distributed storage - where the replicas actually landFailure domains · advanced · ~14 min
- 06Recovery and rebalancing - what a distributed store does after a failureRecovery · advanced · ~15 min
Part LXI
DRBD Concepts
Primary/secondary, replication, split-brain, clustering integration, when DRBD materially helps HA.
- 01DRBD architecture - the distributed replicated block deviceDRBD · advanced · ~10 min
- 02DRBD primary/secondary and split-brain handlingDRBD roles · advanced · ~16 min
- 03DRBD in a cluster - integration with Pacemaker for HADRBD cluster · advanced · ~14 min
- 04DRBD resync and online verification - reading the replication stateDRBD operations · advanced · ~14 min
- 05Dual-primary DRBD - what it actually requiresDRBD dual-primary · expert · ~14 min
- 06When DRBD materially helps HA - and when it does notDRBD decisions · advanced · ~12 min
Part LXII
Cluster Networking
Management, application, storage, heartbeat, OOB, redundant interfaces, switches, VLANs, failure domains.
- 01Cluster network design - management, application, and storageCluster network design · advanced · ~10 min
- 02Management vs application vs storage - the network rolesNetwork roles · intermediate · ~10 min
- 03Redundant interfaces and switches - the cluster network resilienceRedundancy · advanced · ~10 min
- 04VLANs and failure domains in cluster networksVLANs · intermediate · ~10 min
- 05Multicast and unicast on the cluster networkTransport · advanced · ~13 min
- 06Latency, jitter and cluster membershipTiming · expert · ~14 min
Part LXIII
Cluster Time, DNS and Identity Dependencies
DNS unavailable, NTP skew, LDAP unavailable, certificate expiry, designing around dependency failures.
- 01DNS cluster failure impact - what happens when DNS is downDNS failure · intermediate · ~10 min
- 02NTP skew cluster impact - the silent time bombNTP skew · intermediate · ~10 min
- 03LDAP and certificate expiry cluster impact - the silent credential failuresIdentity expiry · intermediate · ~10 min
- 04Mapping cluster dependencies - finding what you depend onDependency mapping · advanced · ~13 min
- 05Designing around dependency failures - remove, cache, degrade, break glassDesign · advanced · ~14 min
- 06Testing dependency failures - injecting the outage safelyTesting · advanced · ~14 min
Part LXIV
Rolling Maintenance
Validate, drain, patch, reboot, validate, return, observe, batch sizing, maintenance mode, health checks.
- 00Rolling maintenance workflow - the loop that keeps a cluster servingWorkflow · advanced · ~12 min
- 01Maintenance mode and drain - the safe cluster updateMaintenance mode · intermediate · ~10 min
- 02Health validation after a change - the safety netHealth validation · intermediate · ~10 min
- 03Batch sizing and rollback - the change control disciplineBatch and rollback · intermediate · ~10 min
- 04Draining behind a load balancer - rolling maintenance without a cluster managerDraining · advanced · ~14 min
- 05Orchestrating the rolling loop - automation that stops on evidenceOrchestration · advanced · ~15 min
Part LXV
Rolling Kernel Upgrades
Kernel package install, reboot requirement, boot validation, rollback, cluster capacity during maintenance, live patching concepts.
- 01Kernel package install - the kernel upgrade workflowKernel install · intermediate · ~10 min
- 02Reboot and boot validation - the kernel upgrade verificationBoot validation · intermediate · ~10 min
- 03Live kernel patching concepts - patching without rebootingLive patching · intermediate · ~10 min
- 04Kernel rollback - getting back to the kernel that workedRollback · advanced · ~12 min
- 05Cluster capacity during a kernel campaign - the reboot budgetCapacity · advanced · ~14 min
- 06Proving a kernel campaign worked - CVEs, backports and vulnerability stateVerification · advanced · ~14 min
Part LXVI
Capacity Planning for Clusters
Normal/peak utilisation, failure capacity, maintenance capacity, growth, N+1 headroom, 90%-on-3-nodes trap.
- 01Capacity utilisation bands - the 90%-on-3-nodes trapUtilisation bands · intermediate · ~10 min
- 02Normal and peak - measuring the load you actually haveMeasurement · intermediate · ~14 min
- 03Utilisation is not saturation - what the queue tells you that the percentage cannotMeasurement · advanced · ~15 min
- 04Forecasting growth - turning a trend into a dateForecasting · intermediate · ~15 min
- 05Maintenance capacity - what a change window costs the clusterFailure and maintenance · advanced · ~14 min
- 06Writing the capacity plan - finding the binding constraintThe plan · advanced · ~15 min
Part LXVII
Cluster Monitoring
Node health, service health, quorum, membership, failovers, resource status, storage, network, replication, time sync.
- 01Cluster monitoring design - what to monitor in a clusterDesign · advanced · ~10 min
- 02Monitoring failovers and resources - the cluster activityFailovers and resources · intermediate · ~14 min
- 03Actionable cluster alerts - the alert that leads to actionActionable alerts · intermediate · ~10 min
- 04Monitoring cluster storage and replication - watching the standby tooStorage and replication · intermediate · ~13 min
- 05Monitoring cluster network and time - the leading indicatorsNetwork and time · advanced · ~13 min
- 06What cluster monitoring cannot tell youLimits · advanced · ~13 min
Part LXVIII
Cluster Incident Response
Ten canonical scenarios: node unreachable, partition, quorum loss, fencing failure, shared storage loss, LB failure, clock skew, DNS outage, deployment regression, memory exhaustion.
- 01Cluster IR: node unreachable - the most common incidentNode unreachable · intermediate · ~10 min
- 02Cluster IR: network partition - deciding which side is realPartition · advanced · ~14 min
- 03Cluster IR: quorum loss - the arithmetic and the decision recordQuorum loss · advanced · ~14 min
- 04Cluster IR: fencing failure and the fence loopFencing failure · advanced · ~15 min
- 05Cluster IR: shared storage loss - when a stop cannot succeedShared storage loss · advanced · ~15 min
- 06Cluster IR: building one timeline from nodes whose clocks disagreeCross-node evidence · advanced · ~15 min
Part LXIX
Hardware Health
SMART, NVMe health, RAID controllers, IPMI, BMC, Redfish concepts, ECC memory, thermal sensors, smartctl, nvme, sensors, ipmitool.
- 01SMART and NVMe health - disk failure predictionSMART NVMe · intermediate · ~10 min
- 02RAID controller monitoring - the storage layer belowRAID controller · advanced · ~10 min
- 03IPMI and BMC basics - the lights-out management foundationIPMI BMC · foundation · ~10 min
- 04Thermal and ECC monitoring - the memory and CPU healthThermal ECC · intermediate · ~10 min
- 05Hardware inventory and the SEL - turning an alert into a part numberInventory and SEL · intermediate · ~13 min
- 06Firmware lifecycle - the layer your package manager cannot seeFirmware · advanced · ~14 min
Part LXX
Out-of-Band Management
IPMI, iDRAC, iLO, Redfish, serial console, remote console, power control, relation to fencing and DR.
- 01OOB architecture - the out-of-band management layerOOB architecture · foundation · ~10 min
- 02Redfish and IPMI APIs - the standards for hardware managementRedfish IPMI · advanced · ~10 min
- 03OOB and DR - the out-of-band layer in disaster recoveryOOB and DR · advanced · ~10 min
- 04Serial console and SOL - the text lifeline into a hostSerial console · advanced · ~14 min
- 05OOB power control - the commands that change statePower control · advanced · ~14 min
- 06OOB credentials and break-glass - access that survives the outageBreak-glass access · advanced · ~14 min
Part LXXI
TLS and PKI
Private keys, certificates, CSRs, CAs, chains, SAN, expiry, revocation, openssl, certificate troubleshooting, expiry incidents.
- 01TLS and PKI concepts - the foundation of encrypted communicationTLS and PKI · foundation · ~13 min
- 02OpenSSL for sysadmins - the practical toolkitOpenSSL · intermediate · ~14 min
- 03Certificate authorities and chains - the trust model in practiceCAs and chains · intermediate · ~16 min
- 04Certificate renewal and expiry - the lifecycle disciplineRenewal expiry · intermediate · ~12 min
- 05TLS troubleshooting - diagnosing failures from the wireTroubleshooting · advanced · ~17 min
- 06Certificate revocation in practice - CRL, OCSP and short lifetimesTroubleshooting · advanced · ~16 min
Part LXXII
Secrets
Passwords, private keys, API tokens, service credentials, secrets in scripts/Git/world-readable/shell history, external managers.
- 01Secrets - the credential management problemProblem · foundation · ~22 min
- 02Secrets at rest - ownership, modes and the copies you forgotAt rest · intermediate · ~16 min
- 03Secrets in scripts, logs and shell historyIn transit through your own tooling · intermediate · ~16 min
- 04Private keys - passphrases, agents and the keys you cannot rotateKey material · intermediate · ~18 min
- 05Rotation - making a credential change a routine operationLifecycle · intermediate · ~17 min
- 06External secret managers - the bootstrap and availability problemsExternal stores · advanced · ~18 min
Part LXXIII
Change Management
Change plans, peer review, pre-checks, rollback, maintenance windows, testing, validation, observation period.
- 01Change plans and peer review - the discipline of safe changesPlans and review · foundation · ~10 min
- 02Change classes and blast radius - sizing the process to the changeClassification · intermediate · ~14 min
- 03Pre-checks and state capture - what to record before you touch anythingExecution · intermediate · ~15 minLab
- 04Staged rollout - canaries, cohorts and the health gateExecution · advanced · ~15 minLab
- 05Abort criteria and the point of no return - rollback as part of the procedureExecution · advanced · ~15 min
- 06Maintenance windows, freezes and the observation periodScheduling · advanced · ~14 min
Part LXXIV
Configuration Drift
How environments become inconsistent, detection, remediation, drift across nodes, packages, firewall rules, configuration.
- 01Drift remediation - bringing hosts back to desired stateRemediation · intermediate · ~10 min
- 02How drift accumulates - the eight ways hosts stop matchingOrigins of drift · intermediate · ~13 min
- 03Detecting drift independently of the CM toolIndependent detection · intermediate · ~15 min
- 04Package and version drift across a fleetPackage drift · intermediate · ~14 min
- 05Runtime versus on-disk drift - when the file is right and the system is notRuntime drift · intermediate · ~15 min
- 06The limits of drift detection - measuring what nothing managesCoverage · advanced · ~15 min
Part LXXV
Immutable vs Mutable Infrastructure
Traditional servers, configuration management, golden images, cloud images, immutable replacement, trade-offs.
- 01Mutable server traditions - the legacy approachMutable server · foundation · ~10 min
- 02Golden images and immutable patterns - the modern approachGolden images · intermediate · ~10 min
- 03Cloud images and the image build pipelineImage pipeline · intermediate · ~17 min
- 04Configuration management and image baking - choosing the seamConfiguration seam · intermediate · ~16 min
- 05State in an immutable fleet - what survives replacement and what does notState · advanced · ~17 min
- 06The trade-offs, and the emergency change that has to existTrade-offs · advanced · ~16 min
Part LXXVI
Virtualisation and Linux
Linux as a VM: virtual hardware, virtio, VMware tools, cloud guest agents, CPU topology, memory ballooning, snapshots, time sync.
- 00Linux as a VM - virtual hardware, guest tools and the clockVirtual hardware · intermediate · ~14 min
- 01Guest agents - the out-of-band channel into a running VMGuest tooling · intermediate · ~16 min
- 02Converting a guest to virtio - the migration that makes a VM unbootableVirtual hardware · intermediate · ~12 min
- 03CPU topology in a guest, and what the guest cannot seeGuest visibility · advanced · ~17 min
- 04Timekeeping in a guest - clock sources, steps and the resumed VMGuest visibility · advanced · ~17 min
- 05Growing a guest online - disk, filesystem, memory and CPUGuest operations · intermediate · ~18 min
Part LXXVII
Linux in the Cloud
cloud-init, metadata services, ephemeral disks, persistent block storage, security groups, IAM concepts, instance lifecycle.
- 00cloud-init and the instance metadata serviceProvisioning · intermediate · ~14 min
- 01Ephemeral and persistent storage - the cloud storage modelEphemeral and persistent · intermediate · ~10 min
- 02Writing user-data - payload formats, module frequency, and the test loopProvisioning · intermediate · ~16 min
- 03Instance lifecycle - reboot, stop/start, terminate, replaceLifecycle · intermediate · ~16 min
- 04Security groups and the host firewall - two layers that fail differentlyNetwork boundary · intermediate · ~16 min
- 05Instance identity and cloud IAM from the Linux sideIdentity · intermediate · ~16 min
Part LXXVIII
Containers from the Linux Perspective
Namespaces, cgroups, OverlayFS, capabilities — the Linux primitives containers build on; cross-link to the Docker course.
- 01Container primitives - namespaces, cgroups, OverlayFS, capabilitiesPrimitives · advanced · ~12 min
- 02OverlayFS for containers - the layered filesystemOverlayFS · intermediate · ~10 min
- 03Network namespaces by hand - building container networking with ipNamespaces · advanced · ~18 min
- 04User namespaces and UID mapping - the kernel side of rootlessNamespaces · advanced · ~18 min
- 05What containers do not isolateBoundaries · advanced · ~17 min
- 06Attributing host symptoms to containers - triage without the runtime CLIOperations · advanced · ~18 min
Part LXXIX
Troubleshooting Methodology
Define symptom, determine impact, check recent changes, collect evidence, identify subsystem, form hypothesis, test safely, restore service, find root cause, prevent recurrence.
- 01The troubleshooting loop - ten steps and their stop-rulesThe loop · intermediate · ~12 minLab
- 02Evidence-based diagnosis - capture before you disturbThe loop · intermediate · ~12 minLab
- 03Preventing recurrence - closing the failure gap and the detection gapThe loop · intermediate · ~10 minLab
- 04Defining the symptom - turning a complaint into a measurementThe loop · intermediate · ~12 min
- 05What changed - working the highest-yield questionThe loop · intermediate · ~14 min
- 06Narrowing to one subsystem - bisect the path, then test one thingThe loop · intermediate · ~14 min
Part LXXX
Common Failure Scenarios
Service won’t start, system won’t boot, filesystem full, inode exhaustion, read-only FS, failed disk, degraded RAID, LVM full, DNS, route, packet loss, firewall, cert expiry, auth, sudo, OOM, CPU saturation, disk latency, broken fstab, kernel upgrade, cluster node loss, quorum loss.
- 01Failure: a service will not start, a host will not bootCommon failures · intermediate · ~12 min
- 02Failure: filesystem full, inodes exhausted, filesystem read-onlyCommon failures · intermediate · ~12 min
- 03Failure: DNS, routing, packet loss, firewall and certificate expiryCommon failures · intermediate · ~14 min
- 04Failure: OOM kills, CPU saturation and disk latencyCommon failures · intermediate · ~14 min
- 05Failure: cluster node loss, quorum loss and dependency outagesCommon failures · advanced · ~14 min
- 06Failure: nobody can log in, and sudo stopped workingCommon failures · intermediate · ~13 min
Part LXXXI
Incident Command
Impact assessment, severity, communication, roles, stabilisation, evidence preservation, recovery, escalation, post-incident review.
- 01Incident roles and communication - the team coordinationRoles and comms · intermediate · ~10 min
- 02Impact assessment - measuring what is broken for whomImpact assessment · intermediate · ~13 min
- 03Stabilisation - restoring service before you understand itStabilisation · intermediate · ~14 min
- 04Evidence preservation - what recovery destroysEvidence preservation · intermediate · ~14 min
- 05Escalation and handover - moving an incident between peopleEscalation · intermediate · ~13 min
- 06Recovery and stand-down - closing an incident properlyRecovery and stand-down · intermediate · ~13 min
Part LXXXII
Root Cause Analysis
Trigger vs contributing factor vs root cause, systemic causes, avoiding "human error" as the explanation.
- 01Avoiding human error as cause - the systemic approachHuman error · foundation · ~10 min
- 02Trigger, contributing factors, root cause - taking an incident apartDecomposition · intermediate · ~14 min
- 03Reconstructing the timeline - evidence over recollectionMethod · advanced · ~16 minLab
- 04Causal chains and stopping rules - not settling for the first plausible answerMethod · advanced · ~16 min
- 05Systemic causes - the conditions that made the incident possibleSystemic · expert · ~16 min
- 06The blameless postmortem - writing an RCA that produces changeOutput · advanced · ~15 min
Part LXXXIII
Operational Documentation
Architecture diagrams, inventories, build procedures, recovery procedures, dependency maps, maintenance procedures, escalation.
- 01Architecture and dependency maps - the documentation foundationArchitecture · foundation · ~10 min
- 02Build and recovery procedures - the operational referenceBuild and recovery · foundation · ~10 min
- 03Runbooks that work at 3amProcedures · intermediate · ~16 min
- 04Inventories, maintenance procedures and escalation pathsRecords · intermediate · ~15 min
- 05The incident record - writing it while it is happeningRecords · intermediate · ~16 min
- 06Documentation that stays true and can be foundMaintenance · intermediate · ~15 min
Part LXXXVI
Capstone: Production Linux Cluster
A 3-node production Linux cluster with HA load balancing, shared storage, monitoring, central logging, configuration management, central identity, DNS, NTP, backup, secrets, certificates — end-to-end operations.
No lessons published in this part yet. The full curriculum is planned in docs/courses/linux/curriculum.md on GitHub.
Part LXXXVII
Break It and Fix It
Major failure-injection section: symptoms first, evidence provided, root cause and remediation hidden behind reveal.