Daily pre-shift check
A 5-minute check performed at the start of each shift.
How to use
Run each command. Note any unexpected output. Investigate anything unusual before the shift progresses.
Items
- cluster-health: Check that the cluster has quorum and Ceph is HEALTH_OK (if in use).
- backup-status: Verify backups ran last night without errors.
- alerts-inbox: Clear or acknowledge any new alerts in your alerting system.
- capacity-warning: Check capacity utilisation; alert if > 80%.
- node-health: Verify all cluster nodes are reachable.
- ongoing-incidents: Review any active incident tickets.
Notes
If anything fails, document in the incident tracker. Don’t fix things immediately unless they are user-impacting.