Free sample question

How would you design Terraform state management across 200+ configurations with locking, isolation, and disaster recovery?

Senior · Remote State · Backends · Chapter 1: State & Backends, question 1 of 7 · from Infrastructure as Code Mastery: Terraform & OpenTofu

What the interviewer is really testing

whether you can reason about blast radius and isolation boundaries at scale, instead of reaching for one giant state file or treating workspaces as real separation.

The 30-second answer

I'd split state by isolation boundary, not by convenience: one backend key per environment per component, so a bad apply to staging networking can never touch prod databases. I use an S3 backend with native state locking (the use_lockfile option, GA in Terraform 1.11, so no separate DynamoDB table), versioning and KMS encryption on the bucket, and a strict naming convention like env/component/terraform.tfstate. Terragrunt keeps that backend config DRY across 200 stacks. DR is just bucket versioning plus cross-region replication, with a tested restore runbook.

Follow-ups the interviewer will probe

How do you recover a corrupted or accidentally deleted state file?
Pull the prior version from S3 versioning into a fresh key, run terraform plan against real infrastructure to confirm it reconciles, and only then promote it. If versioning is gone, rebuild with import blocks. The point is to validate before trusting, and to rehearse this restore before you need it.
A lock is stuck after a crashed pipeline. What now?
Confirm no apply is genuinely still running (check CI, not just the lock), then terraform force-unlock with the lock ID. The senior habit is preventive: short CI timeouts, lock-age alerting, and a documented owner, so a stuck lock is a five-minute fix rather than a team-wide outage.
How do you pass data between separate state files?
Publish stable outputs from the producing stack and consume them via a terraform_remote_state data source, or better, write shared identifiers to SSM Parameter Store so consumers depend on a contract, not on another team's state internals. Keep the dependency graph shallow and one-directional to avoid apply-order deadlocks.

Recall hook

“One state per blast radius.”

split by environment and component so a bad apply stays in its cell; lock, version, and encrypt every cell.

What the book adds to this question

In the ebook every question runs three pages. Between the 30-second answer and the follow-ups it adds a deep dive with a diagram, a decision framework, and a pitfalls-and-signals table. Infrastructure as Code Mastery: Terraform & OpenTofu has 50 questions across 8 chapters and includes the free Interview-Day Playbook.

Want full three-page questions? Download the free 8-question PDF sample (from Cloud Interview Mastery).

Sample questions from the other books