How to use
Quarterly, sitting at the recovery site rather than at a desk in the primary one, with the tier-one service list and the cutover plan open. Most of the seventy minutes goes on four questions the site cannot answer on paper: whether the catalogue there is current, whether the escrow opens, whether every tier-one dependency graph closes without an edge pointing home, and what the last exercise genuinely exercised.
Run the connectivity items from a host at the site, not through a management jump box at the primary. Reading a firewall rule set tells you what somebody wrote; opening the connection tells you what the path does today.
Where an item asks for a date — the escrow retrieval, the login, the capacity comparison, the exercise — a date nobody can produce is a fail. The failure this sheet exists to catch is a site everybody believes in and nobody has used.
Where the numbers come from
Catalogue currency is read at the site: list recovery points from a host there, then compare the newest entry per service with the same query at the primary. One number, taken twice, in two places.
Copy age is measured per source, not per job. What ran at the primary and what arrived at the site are two separate facts, and the recovery reads the second one.
Capacity is compared against observed peak across the last ninety days — CPU, memory, storage, throughput, egress and licence counts — never against the sizing document. Record the comparison date; the figure ages faster than the review interval.
TTL is queried from the authoritative server, not read from the zone file in the repository. Those two disagree from the moment anybody makes a change in a provider console.
The exercise figure is a date plus a scope: tabletop, partial technical recovery of named services, or a cutover carrying real users. A date without a scope is not a number.
Access this needs
Shell access on a host at the recovery site, with the read credential for the catalogue and for the repository copies held there.
The escrow retrieval procedure, and a custodian present or on call for the review. This is the item most often deferred and the one that most often fails when it is finally attempted.
Read access to the tier-one dependency graphs, the cutover plan, capacity figures for both sites, and the certificate inventory for the services the site will present.
Query access to the authoritative DNS servers, and read access to the registrar or provider account. No change authority is required: the only live actions are opening connections that are supposed to be open.
What the review produces
A dated sheet naming the reviewer, the site, the tier-one services in scope and the disposition of every item, carrying five numbers: the age of the catalogue copy at the site, the oldest per-source copy age, the longest customer-facing TTL, the ratio of site capacity to current production peak, and the months since the last exercise.
Attach the dependency-graph findings. Every edge that resolves back to the primary is a defect with a name, an owner and usually a cheap fix, and this is the only review that goes looking for them.
Attach the list of embedded addresses and hostnames that will be wrong at the site. It is the cutover’s longest manual step and the one that can be prepared calmly months ahead.
A failing critical item that is accepted rather than fixed needs a named acceptor and a date. “The escrow has not been opened this year” is a decision somebody is making, and it should be signed.
Sign-off
- Reviewer: ____ Date: ____
- Recovery site owner: ____ Date: ____
- Tier-one service owner: ____ Date: ____
Every critical item must pass. A failing critical item is a blocker rather than a note for the next quarter: record the date, the reviewer, the disposition of every item that did not pass, and the name of whoever accepted the residual risk.