How to use
Out loud, in order, in the last few minutes before the first restore command runs, with the operator and one other person. Every item is a question with an answer that exists or does not exist right now, and none of them takes a minute to produce. Twenty-five minutes covers the whole list even when somebody has to walk away and fetch a passphrase.
This is the moment when the list feels most like an obstacle, and that feeling is the reason it is written down. Somebody senior is asking for an estimate, the service has been down for long enough that people have started apologising to customers, and every item here can be waved through with a plausible sentence. The failures this list catches are not exotic. They are a target that turned out to be the live path, a recovery point that was one hour too new, a restore that had been silently stalled for ninety minutes because nobody had agreed what progress looked like, and a completion declared by the person least able to judge it.
Anything answered with “probably” is answered with “no”. The list has no partial credit, and an item that cannot be closed in a minute is a finding worth the sixty seconds it took to surface, not a reason to move on.
Where the numbers come from
The elapsed-time estimate comes from a restore of this system that somebody performed and timed, retrieval and decryption included, not from the size of the data divided by the speed of the link. If no such number exists, say so on the call and use the largest figure anybody is willing to defend — an admitted guess behaves very differently from a fabricated one when it is exceeded.
Free space comes from the target filesystem in the last few minutes, not from the sizing conversation earlier in the incident. Restored size comes from the metadata written when the backup was taken; a repository is deduplicated and compressed, so its own size understates what it expands to, and it understates it in the direction that fills the target.
The stall threshold is whatever multiple of the estimate the team agreed before starting. Which multiple matters far less than that the number was fixed while nobody had yet invested an hour in the outcome.
Access this needs
Read access to the repository from the recovery host, with the credentials the restore will actually use, plus the passphrase already retrieved. Shell on the target to resolve the path and read free space. The manifest open in a second window. Sight of the runbook, and the incident channel where the estimate, the abort criteria and the named completer are posted.
Nothing on this list needs write access to production. An item that cannot be closed without it has found a second problem, and that problem should be recorded before the restore starts rather than after.
What the review produces
One short block, posted to the incident channel before the first command: the recovery point identifier and the evidence that put it before the damage, whether it was verified or the risk was accepted and by whom, the resolved target, the expected duration and the checkpoint time, the abort criteria, the person who may declare completion and the value they will check, the rollback, and the update cadence.
That block is also the skeleton of the post-incident record. It was written before anybody knew how the night ended, which is the only time these answers can be captured without hindsight quietly improving them.
Sign-off
- Reviewer: ____ Date: ____
- Restore operator: ____ Date: ____
- Authorised to declare completion: ____ Date: ____
Every critical item is closed before the first command runs, or the responder who chose to proceed without it is named in the incident log alongside the item they skipped and the reason.