Infrastructure

Terraform on Azure: diagnose a partial apply before rerunning production

A production runbook for reconstructing an interrupted Terraform apply on Azure, comparing configuration, state and live resources, then choosing rerun, targeted import, rollback or state repair.

12 Aug 2026 azureterraformiacstatedriftci-cdactivity-logresource-graphautomationobservabilityrunbookrollbackproduction

A Terraform pipeline fails halfway through an Azure apply. The job is red, but several operations have already completed: an identity exists, a network rule changed and a role assignment is visible in Azure. Rerunning the pipeline may finish the deployment. It may also collide with a live lock, recreate a resource, hide drift or apply a different plan from the one that was approved.

Terraform does not wrap a complete Azure deployment in one transaction. Azure processes separate operations and Terraform records their results in state as work progresses. After an interruption, Git configuration, remote state and live Azure resources can disagree. This runbook drives one operational decision: resume from a clean plan, import an object created before the failure, repair state under strict control, or roll back the partial change without disturbing healthy resources.

Pin the failed execution

Identify the exact run before touching state. Preserve the commit, workspace, backend key, input variables, Terraform version, provider lockfile, CI identity, approved plan and apply log. A run from another commit or backend is a new change, even if the pipeline name looks identical.

yaml partial-apply-incident.yml
incident: inc-20260812-021
pipeline_run: platform-prod-1842
commit: "<git-sha>"
terraform_version: "<version>"
workspace: prod
backend_key: platform/prod.tfstate
plan_artifact: tfplan-prod-1842
execution_identity: "<federated-principal-id>"
started_at_utc: "<timestamp>"
failed_at_utc: "<timestamp>"
failed_address: "<terraform-resource-address>"

preserve:
- apply log and provider errors
- saved plan and plan JSON
- backend and workspace selection
- state lineage and serial
- Azure activity operations in the same window
- application and platform health evidence

Pause scheduled pipelines against the same state. Confirm that no apply, drift job or operator workstation is still writing. Remove a stale lock only after proving that its owner has stopped; otherwise two writers can turn an interrupted apply into state corruption.

Capture state before refresh

Pull a raw copy of remote state before running commands that refresh objects. The pre-refresh view is evidence of what Terraform knew when the run stopped.

bash 01-capture-terraform-state.sh
set -euo pipefail

terraform version
terraform workspace show
terraform providers lock -platform=linux_amd64

terraform state pull > state-before.json
jq '{lineage, serial, terraform_version, resources: [.resources[].type]}' state-before.json > state-before-summary.json

terraform show -json > state-view-before.json
terraform state list > state-addresses-before.txt

The lineage must identify the expected state and the serial must remain stable while writers are frozen. If it moves, stop local recovery and find the process that wrote a new version.

Treat state as sensitive material. It can contain identifiers, connection strings and secret values. Keep captures in an approved incident evidence location, restrict access and apply the normal retention policy.

Rebuild the Azure side of the apply

The final pipeline message is often the first visible error, not the last successful operation. Azure Activity Log shows ARM writes, status, caller and correlation IDs. Azure Resource Graph can then confirm the current shape of affected resources.

bash 02-azure-activity-window.sh
SUBSCRIPTION_ID="<subscription-id>"
START_UTC="<apply-start-utc>"
END_UTC="<apply-end-utc>"

az account set --subscription "$SUBSCRIPTION_ID"

az monitor activity-log list --start-time "$START_UTC" --end-time "$END_UTC" --status Succeeded --query "[].{time:eventTimestamp,operation:operationName.value,resource:resourceId,caller:caller,correlationId:correlationId}" --output json > azure-writes-succeeded.json

az monitor activity-log list --start-time "$START_UTC" --end-time "$END_UTC" --status Failed --output json > azure-writes-failed.json

Narrow the result to the intended resource groups and types. An ARM Succeeded event proves that Azure accepted an operation. It does not prove that Terraform persisted the object or that the consuming service is healthy. Likewise, an object that exists in Azure but not in state is a recovery candidate, not an automatic import.

Compare configuration, state and reality

Build a three-way record for each Terraform address touched by the saved plan. This prevents the incident from being labelled “broken state” when the actual problem is the wrong workspace, a successful Azure creation followed by a provider timeout, policy enforcement or out-of-band drift.

text three-way-state-matrix.txt
For every affected address
Configuration
  resource exists in the executed commit
  address is stable or covered by a moved block
  expected arguments and dependencies are known

Remote state
  address is present or absent
  recorded Azure resource ID
  attributes known after the last refresh
  expected lineage and serial

Live Azure
  resource is present or absent
  current properties and owning identity
  create or update event in Activity Log
  observable service impact

Classification
  all three views converge
  Azure object exists, state address missing
  state address exists, Azure object missing
  state and Azure exist with different properties
  wrong backend or workspace
  ambiguous IaC ownership

Start with resources in the approved plan and their direct dependencies. A subscription-wide hunt adds noise and weakens causality during recovery.

Use refresh-only as evidence

After preserving state and excluding other writers, produce a refresh-only plan. It shows how Terraform would reconcile its recorded view with remote objects without mixing in normal configuration changes.

bash 03-refresh-only-evidence.sh
terraform plan -refresh-only -out=tfrefresh

terraform show -json tfrefresh > tfrefresh.json

jq -r '
.resource_changes[]?
| select(.change.actions != ["no-op"])
| [.address, (.change.actions | join(","))]
| @tsv
' tfrefresh.json > refresh-differences.tsv

Do not apply that plan automatically. It is diagnostic evidence first. If refresh proposes removing a critical object from state because Azure no longer returns it, verify subscription, tenant, permissions and resource ID before allowing a write.

Next, create a normal plan from the exact failed commit. A rerun is blocked when it would recreate a live object, destroy a healthy resource or change anything outside the reviewed scope.

Classify where execution broke

The recovery path depends on the boundary where state and reality separated.

text partial-apply-classification.txt
Failure before an Azure call
No matching write in Activity Log
State and Azure remain coherent
Fix input, permission or provider error and generate a new plan

Azure created the object, Terraform did not record it
ARM operation succeeded, object is live, state address is absent
Confirm ownership and consider a targeted import

State updated, pipeline later failed
Address and Azure ID are coherent, Azure write succeeded
The next plan may be empty for this resource
Diagnose the next failed dependency instead of recreating it

Dependency or policy denied
Earlier resources are healthy, the next operation was rejected
Correct the cause and review the complete plan again

Wrong backend or workspace
Lineage, serial or inventory is unexpected
Import nothing; restore the correct execution context first

This classification prevents a false rollback that deletes a correctly created resource only because the overall job is red. It also prevents a false recovery that leaves a live Azure object outside state governance.

Repair ownership before state

Import is appropriate when the live resource matches approved configuration, belongs to this Terraform stack and has one unambiguous address. Prefer a versioned import block when the repository workflow supports it.

hcl recovery-import.tf
import {
to = azurerm_resource_group.example
id = "/subscriptions/<subscription-id>/resourceGroups/<resource-group>"
}

Review a full plan after import. An immediate large update or replacement means the configuration does not accurately describe the object you just adopted.

Reserve terraform state rm, state mv and direct state replacement for cases where ownership is documented, a backup is verified and a second reviewer is present. Removing an address from state does not delete its Azure object; it removes Terraform governance. Manually pushing modified state is a last-resort operation that must preserve the expected lineage, advance the serial correctly and have a tested restoration path.

Decide rerun, import, rollback or stop

Record the recovery gate in language that another operator can review.

yaml partial-apply-decision.yml
rerun:
when:
  - no active writer remains
  - backend, workspace, commit and provider lock match
  - state and Azure ownership are coherent
  - a newly reviewed full plan contains only expected actions

targeted_import:
when:
  - Azure creation succeeded before the interruption
  - the resource matches approved configuration
  - its Terraform address and owner are unambiguous
  - the post-import plan is reviewed

rollback:
when:
  - a completed partial change causes proven impact
  - the reverse plan is smaller and understood
  - application and infrastructure validation are ready

stop_and_escalate:
when:
  - state lineage or serial changed unexpectedly
  - another writer cannot be excluded
  - import would seize an ambiguously owned resource
  - the recovery plan deletes or replaces an unrelated resource

-target is not a normal deployment mechanism. It can support an exceptional, documented recovery when one dependency must be repaired first, but a complete plan must follow. Otherwise the team may close the incident while the rest of the graph remains divergent.

Validate convergence and the rollback path

After repair, create a full plan with the same backend and input set. The expected result is not necessarily an empty plan; it is a plan that matches the approved recovery decision exactly. Apply it with one writer, a working lock and retained logs.

Validate the resources and the service consuming them: effective identity, expected route or rule, dependency access, logs, metrics and a correlated synthetic probe. Confirm that remote state advanced by the expected serial and that a new plan exposes no unexplained residual changes.

Rollback from a partial apply also needs a reviewed plan. Returning to the previous commit may propose deleting resources created during the interrupted run. Treat that reverse plan as a production change: protect data, validate impact and preserve healthy resources when deletion would make the incident worse.

Conclusion

A red Terraform job does not mean that nothing changed or that everything should be undone. After a partial apply, the useful source of truth is a controlled comparison between the executed commit, remote state and operations Azure actually accepted.

Recovery is defensible when the backend is pinned, concurrent writers are excluded, every resource has a clear owner and the new plan is narrower than the incident. Rerun when all three views converge, import when Azure created an object that clearly belongs to the stack, roll back when a partial change causes proven harm, and stop all state intervention while lineage, serial or ownership remain ambiguous.