How to use this checklist
The control plane is the only component whose failure you cannot route
around with more replicas of the workload. Work this list on a cluster you
believe is already healthy — the point is to find the single points of
failure that a green kubectl get nodes hides.
Mark an item N/A when it genuinely does not apply, and write down why. An unexplained N/A is the most common way a checklist stops working.
Sign-off
Every critical item must pass. A failing critical item blocks the deployment or the maintenance window; it is not a note for later. Record the date, the reviewer, and the disposition of every item that did not pass.