TerraformXV · Environment Architecture and State BoundariesProduction Terraform
State Boundaries as the Unit of Failure
What you'll learn
- Design state boundaries
- Identify the failure domain
- Plan the boundaries
Prerequisites
None — start here.
Verified against Terraform CLI 1.9.x · OpenTofu 1.7.x · HCL 2.0 · bpg/proxmox provider 0.66+ · hashicorp/local provider 2.5+ · hashicorp/null provider 3.2+ · hashicorp/random provider 3.6+ · hashicorp/http provider 3.4+ · Ubuntu 24.04 LTS · Debian 12 (Bookworm)
Objective
Design state boundaries
What this lesson covers
- Identify the failure domain
- Plan the boundaries
Why this matters in production
Production Terraform operations have blast radius. State Boundaries as the Unit of Failure is one of the operational controls that determines whether a change is safe to apply. Skip this lesson at the cost of not understanding a production control.
How it works
Provide an overview of the topic. This lesson covers the state boundaries as the unit of failure concept, the underlying mechanism, and the production implications.
Key concepts.
- The terminology used in the topic.
- The mechanism that determines the operational behaviour.
- The production control that mitigates the operational risk.
Operational implications.
The state boundaries as the unit of failure concept affects the production change-management workflow. A change in this area has blast radius across the entire estate.
How to configure it
Apply the configuration pattern:
- Set up the configuration block.
- Validate the configuration.
- Run the plan.
- Verify the operations.
- Document the change.
How to inspect it
Inspection pattern:
- Use the relevant CLI command to inspect the state.
- Verify the configuration matches the expectation.
- Confirm the plan is empty.
Validation
The validation pattern:
- Confirm the configuration is valid.
- Verify the plan is empty.
- Confirm the operations are correct.
Production failure modes
The following failure modes are the most common in production:
- Misconfiguration of the relevant parameter.
- Drift between the configuration and the real world.
- Provider failure during the apply.
- State corruption or loss.
How to recover
The recovery procedure:
- Investigate the failure.
- Identify the cause.
- Apply the remediation.
- Verify the state.
- Document the incident.
Related runbooks
The following runbooks in this course apply to this lesson:
- terraform-runbook-investigate-state-lock
- terraform-runbook-recover-partial-apply
- terraform-runbook-investigate-provider-failure
What comes next
The next lesson in XV-Environment-Architecture continues the exploration of State Boundaries as the Unit of Failure. Follow the order to maintain the production-first narrative.
Knowledge check · 7 questions
Q1. Why are multiple environments important?
Q2. What is a state boundary?
Q3. Workspaces are appropriate for production isolation.
Q4. What is the role of directories in multi-environment estates?
Q5. Which of the following are good production patterns for environments? (Select all that apply.)
Q6. What is the role of accounts/projects/subscriptions in environments?
Q7. A team uses workspaces for staging and production. The state is corrupted in staging. Production is unaffected. What is the fix?
Passing score: 75%. Answers are checked in this browser.