TerraformXXVIII · Disaster Recovery and ResilienceProduction Terraform
Periodic DR Testing
What you'll learn
- Schedule DR tests
- Validate the recovery
- Document the tests
Prerequisites
None — start here.
Verified against Terraform CLI 1.9.x · OpenTofu 1.7.x · HCL 2.0 · bpg/proxmox provider 0.66+ · hashicorp/local provider 2.5+ · hashicorp/null provider 3.2+ · hashicorp/random provider 3.6+ · hashicorp/http provider 3.4+ · Ubuntu 24.04 LTS · Debian 12 (Bookworm)
Objective
Schedule DR tests
What this lesson covers
- Validate the recovery
- Document the tests
Why this matters in production
Production Terraform operations have blast radius. Periodic DR Testing is one of the operational controls that determines whether a change is safe to apply. Skip this lesson at the cost of not understanding a production control.
How it works
Provide an overview of the topic. This lesson covers the periodic dr testing concept, the underlying mechanism, and the production implications.
Key concepts.
- The terminology used in the topic.
- The mechanism that determines the operational behaviour.
- The production control that mitigates the operational risk.
Operational implications.
The periodic dr testing concept affects the production change-management workflow. A change in this area has blast radius across the entire estate.
How to configure it
Apply the configuration pattern:
- Set up the configuration block.
- Validate the configuration.
- Run the plan.
- Verify the operations.
- Document the change.
How to inspect it
Inspection pattern:
- Use the relevant CLI command to inspect the state.
- Verify the configuration matches the expectation.
- Confirm the plan is empty.
Validation
The validation pattern:
- Confirm the configuration is valid.
- Verify the plan is empty.
- Confirm the operations are correct.
Production failure modes
The following failure modes are the most common in production:
- Misconfiguration of the relevant parameter.
- Drift between the configuration and the real world.
- Provider failure during the apply.
- State corruption or loss.
How to recover
The recovery procedure:
- Investigate the failure.
- Identify the cause.
- Apply the remediation.
- Verify the state.
- Document the incident.
Related runbooks
The following runbooks in this course apply to this lesson:
- terraform-runbook-investigate-state-lock
- terraform-runbook-recover-partial-apply
- terraform-runbook-investigate-provider-failure
What comes next
The next lesson in XXVIII-Disaster-Recovery continues the exploration of Periodic DR Testing. Follow the order to maintain the production-first narrative.
Knowledge check · 7 questions
Q1. What is RPO in Terraform DR?
Q2. What is RTO in Terraform DR?
Q3. Cross-region state replication is appropriate for production Terraform DR.
Q4. What is the role of the execution environment in DR?
Q5. Which of the following are required for a production DR plan? (Select all that apply.)
Q6. What is the role of state versioning in DR?
Q7. A state backend fails. The team needs to recover. The first step is to:
Passing score: 75%. Answers are checked in this browser.