Weekly operational review
A 30-60 minute review of cluster health and operational metrics. Performed by the on-call engineer or a designated reviewer.
Goals
- Catch small problems before they become incidents
- Validate backup and DR procedures
- Identify trends that need attention (capacity, performance)
What to look for
- Backup success rate: anything below 95% needs investigation
- Disk space: anything above 80% needs a forecast; above 90% is urgent
- Pending updates: prioritise security, schedule restarts
- Failed login patterns: spikes may indicate attack or misconfiguration