Infrastructure
Terraform on Azure: diagnose a partial apply before rerunning production
A production runbook for reconstructing an interrupted Terraform apply on Azure, comparing configuration, state and live resources, then choosing rerun, targeted import, rollback or state repair.
A Terraform pipeline fails halfway through an Azure apply. The job is red, but several operations have already completed: an identity exists, a network rule changed and a role assignment is visible in Azure. Rerunning the pipeline may finish the deployment. It may also collide with a live lock, recreate a resource, hide drift or apply a different plan from the one that was approved.
Terraform does not wrap a complete Azure deployment in one transaction. Azure processes separate operations and Terraform records their results in state as work progresses. After an interruption, Git configuration, remote state and live Azure resources can disagree. This runbook drives one operational decision: resume from a clean plan, import an object created before the failure, repair state under strict control, or roll back the partial change without disturbing healthy resources.
Pin the failed execution
Identify the exact run before touching state. Preserve the commit, workspace, backend key, input variables, Terraform version, provider lockfile, CI identity, approved plan and apply log. A run from another commit or backend is a new change, even if the pipeline name looks identical.
incident: inc-20260812-021
pipeline_run: platform-prod-1842
commit: "<git-sha>"
terraform_version: "<version>"
workspace: prod
backend_key: platform/prod.tfstate
plan_artifact: tfplan-prod-1842
execution_identity: "<federated-principal-id>"
started_at_utc: "<timestamp>"
failed_at_utc: "<timestamp>"
failed_address: "<terraform-resource-address>"
preserve:
- apply log and provider errors
- saved plan and plan JSON
- backend and workspace selection
- state lineage and serial
- Azure activity operations in the same window
- application and platform health evidence Pause scheduled pipelines against the same state. Confirm that no apply, drift job or operator workstation is still writing. Remove a stale lock only after proving that its owner has stopped; otherwise two writers can turn an interrupted apply into state corruption.
Capture state before refresh
Pull a raw copy of remote state before running commands that refresh objects. The pre-refresh view is evidence of what Terraform knew when the run stopped.
set -euo pipefail
terraform version
terraform workspace show
terraform providers lock -platform=linux_amd64
terraform state pull > state-before.json
jq '{lineage, serial, terraform_version, resources: [.resources[].type]}' state-before.json > state-before-summary.json
terraform show -json > state-view-before.json
terraform state list > state-addresses-before.txt The lineage must identify the expected state and the serial must remain stable while writers are frozen. If it moves, stop local recovery and find the process that wrote a new version.
Treat state as sensitive material. It can contain identifiers, connection strings and secret values. Keep captures in an approved incident evidence location, restrict access and apply the normal retention policy.
Rebuild the Azure side of the apply
The final pipeline message is often the first visible error, not the last successful operation. Azure Activity Log shows ARM writes, status, caller and correlation IDs. Azure Resource Graph can then confirm the current shape of affected resources.
SUBSCRIPTION_ID="<subscription-id>"
START_UTC="<apply-start-utc>"
END_UTC="<apply-end-utc>"
az account set --subscription "$SUBSCRIPTION_ID"
az monitor activity-log list --start-time "$START_UTC" --end-time "$END_UTC" --status Succeeded --query "[].{time:eventTimestamp,operation:operationName.value,resource:resourceId,caller:caller,correlationId:correlationId}" --output json > azure-writes-succeeded.json
az monitor activity-log list --start-time "$START_UTC" --end-time "$END_UTC" --status Failed --output json > azure-writes-failed.json Narrow the result to the intended resource groups and types. An ARM Succeeded event proves that Azure accepted an operation. It does not prove that Terraform persisted the object or that the consuming service is healthy. Likewise, an object that exists in Azure but not in state is a recovery candidate, not an automatic import.
Compare configuration, state and reality
Build a three-way record for each Terraform address touched by the saved plan. This prevents the incident from being labelled “broken state” when the actual problem is the wrong workspace, a successful Azure creation followed by a provider timeout, policy enforcement or out-of-band drift.
For every affected address
Configuration
resource exists in the executed commit
address is stable or covered by a moved block
expected arguments and dependencies are known
Remote state
address is present or absent
recorded Azure resource ID
attributes known after the last refresh
expected lineage and serial
Live Azure
resource is present or absent
current properties and owning identity
create or update event in Activity Log
observable service impact
Classification
all three views converge
Azure object exists, state address missing
state address exists, Azure object missing
state and Azure exist with different properties
wrong backend or workspace
ambiguous IaC ownership Start with resources in the approved plan and their direct dependencies. A subscription-wide hunt adds noise and weakens causality during recovery.
Use refresh-only as evidence
After preserving state and excluding other writers, produce a refresh-only plan. It shows how Terraform would reconcile its recorded view with remote objects without mixing in normal configuration changes.
terraform plan -refresh-only -out=tfrefresh
terraform show -json tfrefresh > tfrefresh.json
jq -r '
.resource_changes[]?
| select(.change.actions != ["no-op"])
| [.address, (.change.actions | join(","))]
| @tsv
' tfrefresh.json > refresh-differences.tsv Do not apply that plan automatically. It is diagnostic evidence first. If refresh proposes removing a critical object from state because Azure no longer returns it, verify subscription, tenant, permissions and resource ID before allowing a write.
Next, create a normal plan from the exact failed commit. A rerun is blocked when it would recreate a live object, destroy a healthy resource or change anything outside the reviewed scope.
Classify where execution broke
The recovery path depends on the boundary where state and reality separated.
Failure before an Azure call
No matching write in Activity Log
State and Azure remain coherent
Fix input, permission or provider error and generate a new plan
Azure created the object, Terraform did not record it
ARM operation succeeded, object is live, state address is absent
Confirm ownership and consider a targeted import
State updated, pipeline later failed
Address and Azure ID are coherent, Azure write succeeded
The next plan may be empty for this resource
Diagnose the next failed dependency instead of recreating it
Dependency or policy denied
Earlier resources are healthy, the next operation was rejected
Correct the cause and review the complete plan again
Wrong backend or workspace
Lineage, serial or inventory is unexpected
Import nothing; restore the correct execution context first This classification prevents a false rollback that deletes a correctly created resource only because the overall job is red. It also prevents a false recovery that leaves a live Azure object outside state governance.
Repair ownership before state
Import is appropriate when the live resource matches approved configuration, belongs to this Terraform stack and has one unambiguous address. Prefer a versioned import block when the repository workflow supports it.
import {
to = azurerm_resource_group.example
id = "/subscriptions/<subscription-id>/resourceGroups/<resource-group>"
} Review a full plan after import. An immediate large update or replacement means the configuration does not accurately describe the object you just adopted.
Reserve terraform state rm, state mv and direct state replacement for cases where ownership is documented, a backup is verified and a second reviewer is present. Removing an address from state does not delete its Azure object; it removes Terraform governance. Manually pushing modified state is a last-resort operation that must preserve the expected lineage, advance the serial correctly and have a tested restoration path.
Decide rerun, import, rollback or stop
Record the recovery gate in language that another operator can review.
rerun:
when:
- no active writer remains
- backend, workspace, commit and provider lock match
- state and Azure ownership are coherent
- a newly reviewed full plan contains only expected actions
targeted_import:
when:
- Azure creation succeeded before the interruption
- the resource matches approved configuration
- its Terraform address and owner are unambiguous
- the post-import plan is reviewed
rollback:
when:
- a completed partial change causes proven impact
- the reverse plan is smaller and understood
- application and infrastructure validation are ready
stop_and_escalate:
when:
- state lineage or serial changed unexpectedly
- another writer cannot be excluded
- import would seize an ambiguously owned resource
- the recovery plan deletes or replaces an unrelated resource -target is not a normal deployment mechanism. It can support an exceptional, documented recovery when one dependency must be repaired first, but a complete plan must follow. Otherwise the team may close the incident while the rest of the graph remains divergent.
Validate convergence and the rollback path
After repair, create a full plan with the same backend and input set. The expected result is not necessarily an empty plan; it is a plan that matches the approved recovery decision exactly. Apply it with one writer, a working lock and retained logs.
Validate the resources and the service consuming them: effective identity, expected route or rule, dependency access, logs, metrics and a correlated synthetic probe. Confirm that remote state advanced by the expected serial and that a new plan exposes no unexplained residual changes.
Rollback from a partial apply also needs a reviewed plan. Returning to the previous commit may propose deleting resources created during the interrupted run. Treat that reverse plan as a production change: protect data, validate impact and preserve healthy resources when deletion would make the incident worse.
Conclusion
A red Terraform job does not mean that nothing changed or that everything should be undone. After a partial apply, the useful source of truth is a controlled comparison between the executed commit, remote state and operations Azure actually accepted.
Recovery is defensible when the backend is pinned, concurrent writers are excluded, every resource has a clear owner and the new plan is narrower than the incident. Rerun when all three views converge, import when Azure created an object that clearly belongs to the stack, roll back when a partial change causes proven harm, and stop all state intervention while lineage, serial or ownership remain ambiguous.