Skip to main content
RunBook Academy

TerraformXXIX · Incident Response: The 3 AM TestProduction Terraform

Preventive Measures from Incidents

Intermediate⏱ ~12 min

What you'll learn

  • Identify preventive measures
  • Apply the changes
  • Track the follow-up

Prerequisites

None — start here.

Verified against Terraform CLI 1.9.x · OpenTofu 1.7.x · HCL 2.0 · bpg/proxmox provider 0.66+ · hashicorp/local provider 2.5+ · hashicorp/null provider 3.2+ · hashicorp/random provider 3.6+ · hashicorp/http provider 3.4+ · Ubuntu 24.04 LTS · Debian 12 (Bookworm)

Not yet marked complete on this device.

Objective

Identify preventive measures

What this lesson covers

  • Apply the changes
  • Track the follow-up

Why this matters in production

Production Terraform operations have blast radius. Preventive Measures from Incidents is one of the operational controls that determines whether a change is safe to apply. Skip this lesson at the cost of not understanding a production control.

How it works

Provide an overview of the topic. This lesson covers the preventive measures from incidents concept, the underlying mechanism, and the production implications.

Key concepts.

  • The terminology used in the topic.
  • The mechanism that determines the operational behaviour.
  • The production control that mitigates the operational risk.

Operational implications.

The preventive measures from incidents concept affects the production change-management workflow. A change in this area has blast radius across the entire estate.

How to configure it

Apply the configuration pattern:

  1. Set up the configuration block.
  2. Validate the configuration.
  3. Run the plan.
  4. Verify the operations.
  5. Document the change.

How to inspect it

Inspection pattern:

  • Use the relevant CLI command to inspect the state.
  • Verify the configuration matches the expectation.
  • Confirm the plan is empty.

Validation

The validation pattern:

  • Confirm the configuration is valid.
  • Verify the plan is empty.
  • Confirm the operations are correct.

Production failure modes

The following failure modes are the most common in production:

  1. Misconfiguration of the relevant parameter.
  2. Drift between the configuration and the real world.
  3. Provider failure during the apply.
  4. State corruption or loss.

How to recover

The recovery procedure:

  1. Investigate the failure.
  2. Identify the cause.
  3. Apply the remediation.
  4. Verify the state.
  5. Document the incident.

The following runbooks in this course apply to this lesson:

  • terraform-runbook-investigate-state-lock
  • terraform-runbook-recover-partial-apply
  • terraform-runbook-investigate-provider-failure

What comes next

The next lesson in XXIX-Incident-Response continues the exploration of Preventive Measures from Incidents. Follow the order to maintain the production-first narrative.

Knowledge check · 7 questions

  1. Q1. What is the first step in incident response?

  2. Q2. What is break-glass procedure?

  3. Q3. Post-incident review is optional.

  4. Q4. What is the role of preventive measures in incident response?

  5. Q5. Which of the following are part of the 3 AM test? (Select all that apply.)

  6. Q6. What is the role of runbooks in incident response?

  7. Q7. A production apply fails at 3 AM. The team is in incident mode. The first step is to:

Passing score: 75%. Answers are checked in this browser.