TerraformXXIX · Incident Response: The 3 AM TestProduction Terraform
Preventive Measures from Incidents
What you'll learn
- Identify preventive measures
- Apply the changes
- Track the follow-up
Prerequisites
None — start here.
Verified against Terraform CLI 1.9.x · OpenTofu 1.7.x · HCL 2.0 · bpg/proxmox provider 0.66+ · hashicorp/local provider 2.5+ · hashicorp/null provider 3.2+ · hashicorp/random provider 3.6+ · hashicorp/http provider 3.4+ · Ubuntu 24.04 LTS · Debian 12 (Bookworm)
Objective
Identify preventive measures
What this lesson covers
- Apply the changes
- Track the follow-up
Why this matters in production
Production Terraform operations have blast radius. Preventive Measures from Incidents is one of the operational controls that determines whether a change is safe to apply. Skip this lesson at the cost of not understanding a production control.
How it works
Provide an overview of the topic. This lesson covers the preventive measures from incidents concept, the underlying mechanism, and the production implications.
Key concepts.
- The terminology used in the topic.
- The mechanism that determines the operational behaviour.
- The production control that mitigates the operational risk.
Operational implications.
The preventive measures from incidents concept affects the production change-management workflow. A change in this area has blast radius across the entire estate.
How to configure it
Apply the configuration pattern:
- Set up the configuration block.
- Validate the configuration.
- Run the plan.
- Verify the operations.
- Document the change.
How to inspect it
Inspection pattern:
- Use the relevant CLI command to inspect the state.
- Verify the configuration matches the expectation.
- Confirm the plan is empty.
Validation
The validation pattern:
- Confirm the configuration is valid.
- Verify the plan is empty.
- Confirm the operations are correct.
Production failure modes
The following failure modes are the most common in production:
- Misconfiguration of the relevant parameter.
- Drift between the configuration and the real world.
- Provider failure during the apply.
- State corruption or loss.
How to recover
The recovery procedure:
- Investigate the failure.
- Identify the cause.
- Apply the remediation.
- Verify the state.
- Document the incident.
Related runbooks
The following runbooks in this course apply to this lesson:
- terraform-runbook-investigate-state-lock
- terraform-runbook-recover-partial-apply
- terraform-runbook-investigate-provider-failure
What comes next
The next lesson in XXIX-Incident-Response continues the exploration of Preventive Measures from Incidents. Follow the order to maintain the production-first narrative.
Knowledge check · 7 questions
Q1. What is the first step in incident response?
Q2. What is break-glass procedure?
Q3. Post-incident review is optional.
Q4. What is the role of preventive measures in incident response?
Q5. Which of the following are part of the 3 AM test? (Select all that apply.)
Q6. What is the role of runbooks in incident response?
Q7. A production apply fails at 3 AM. The team is in incident mode. The first step is to:
Passing score: 75%. Answers are checked in this browser.