Automation
Terraform on Azure: contain concurrent applies before state divergence
A production runbook for proving which Terraform writer owns an Azure Blob state lease, stopping competing pipelines, recovering a stale lock and validating state before another apply.
Two Terraform pipelines start against the same Azure environment. The first is applying a network change; the second waits for the remote state lock, then an operator considers force-unlock to get delivery moving again. If the first writer is still alive, removing its lock does not fix the queue. It allows two processes to make decisions from different views of the same infrastructure.
The use case is an azurerm backend stored as a blob in Azure Storage, shared by scheduled drift checks, pull-request plans and production applies. This runbook identifies the exact state, proves whether a writer is active, contains competing jobs, recovers only a stale lock, and validates convergence before another apply. The objective is not to clear the lock quickly. It is to restore one authoritative writer.
Freeze the state scope
Start with the backend identity, not the pipeline display name. Two jobs can look related while using different keys or workspaces. Conversely, jobs from different repositories can collide on the same blob.
incident: inc-20260826-terraform-lock
environment: production
backend:
storage_account: sttfstateprodweu
container: tfstate
key: platform/prod.tfstate
workspace: default
contenders:
- pipeline: platform-apply-1849
commit: <git-sha>
operation: apply
started_utc: <timestamp>
- pipeline: nightly-drift-921
commit: <git-sha>
operation: plan
started_utc: <timestamp>
stop_conditions:
- another writer cannot be excluded
- state serial changes during diagnosis
- backend key or workspace is ambiguous
- live Azure writes continue after jobs are cancelled Pause every scheduled or manually triggered job that targets this state. Disable automatic retries temporarily. Do not use -lock=false; the lock is the control preventing a second writer from entering the incident.
Preserve the lock diagnostic
Terraform lock errors normally expose a lock ID plus information such as path, operation, owner, version and creation time. Keep the complete error with the pipeline run. The lock ID is not an arbitrary confirmation token: terraform force-unlock uses it to target the lock Terraform reported.
Evidence to preserve
Lock ID and backend path
Operation: plan, apply or another state write
Owner or runner identity
Terraform version
Lock creation time
Pipeline run IDs and current status
Last log timestamp from each runner
Commit, workspace and backend configuration
Azure Activity Log window for the apply identity Do not infer that a lock is stale only from its age. A long apply, a slow provider operation or a runner that lost log streaming can still be active. Age is a lead; writer liveness is the decision criterion.
Prove whether the writer is still active
Correlate three views: the CI control plane, the runner process and Azure operations. A job marked cancelled may still have a child process running. A runner that disappeared may have sent an Azure request immediately before losing connectivity.
SUBSCRIPTION_ID="<subscription-id>"
START_UTC="<lock-created-utc>"
END_UTC="<now-utc>"
az account set --subscription "$SUBSCRIPTION_ID"
az monitor activity-log list --start-time "$START_UTC" --end-time "$END_UTC" --query "[].{time:eventTimestamp,status:status.value,operation:operationName.value,resource:resourceId,caller:caller,correlationId:correlationId}" --output json > azure-operations-during-lock.json Filter the result to the execution identity and expected scopes. Recent writes do not prove that Terraform still owns a process, but they prevent a premature unlock until the related job and operation are understood.
Classify the incident before acting:
- Active writer: the process is running or Azure operations continue. Wait or stop that writer cleanly.
- Stale Terraform lock: the owning process is gone, no retry is active and the state serial is stable. A controlled
force-unlockis possible. - Wrong execution context: the reported path, workspace or tenant is not the intended one. Correct the context; unlock nothing.
- Backend access failure: authentication, DNS or Storage reachability failed before lock acquisition. Repair that path instead of treating it as contention.
Read the Azure Blob lease as supporting evidence
The azurerm backend uses Azure Blob Storage native locking. Blob properties can confirm that the expected state object has an active lease and show its last modification. They do not identify the live Terraform process by themselves.
ACCOUNT="sttfstateprodweu"
CONTAINER="tfstate"
BLOB="platform/prod.tfstate"
az storage blob show --auth-mode login --account-name "$ACCOUNT" --container-name "$CONTAINER" --name "$BLOB" --query "{etag:properties.etag,lastModified:properties.lastModified,leaseStatus:properties.lease.status,leaseState:properties.lease.state,leaseDuration:properties.lease.duration}" --output json Use a read-only identity for this inspection when possible. Directly breaking the blob lease bypasses Terraform’s lock workflow and can remove protection from a valid writer. Keep Azure lease-break operations for an exceptional recovery procedure with a storage and Terraform state owner present, not as the normal unlock path.
Stop contenders before removing a stale lock
Cancel queued jobs first, then terminate the owning runner or process through the CI platform. Confirm that it cannot be restarted automatically. Observe a short, explicit quiet window: no new runner heartbeat, no Azure writes by the execution identity and no state modification.
Record the state metadata without exposing its contents broadly.
set -euo pipefail
umask 077
terraform version
terraform workspace show
terraform state pull > state-before-unlock.json
jq '{lineage,serial,terraform_version}' state-before-unlock.json > state-before-unlock-metadata.json
sha256sum state-before-unlock.json > state-before-unlock.sha256 The state file may contain sensitive values. Store it only in an approved evidence location, restrict access and delete the working copy according to the incident procedure.
If the owner is gone, the backend identity is exact and the serial remains stable, use the lock ID reported by Terraform.
LOCK_ID="<terraform-lock-id>"
terraform force-unlock "$LOCK_ID" The command removes the lock; it does not undo Azure resources or repair state. If it rejects the ID, stop and refresh the evidence. Do not substitute a different lock or break the lease just to make the command pass.
Re-enter through a clean plan
Allow one diagnostic job only. Use the same commit, Terraform version, provider lockfile, variables, workspace and backend configuration as the interrupted run. Acquire the normal lock with a bounded wait and produce a new plan.
set -euo pipefail
terraform init -input=false
terraform workspace show
terraform plan -input=false -lock-timeout=5m -out=recovery.tfplan
terraform show -json recovery.tfplan > recovery.tfplan.json
jq -r '
.resource_changes[]?
| select(.change.actions != ["no-op"])
| [.address, (.change.actions | join(","))]
| @tsv
' recovery.tfplan.json > recovery-actions.tsv An empty or expected plan supports recovery. A plan that deletes, replaces or imports unrelated resources means the incident has moved beyond a stale lock. Compare configuration, remote state and live Azure resources before applying anything.
Decide wait, unlock, recover or stop
Use a decision that another operator can challenge:
WAIT
The owning writer is active or its Azure operation is unresolved.
FORCE-UNLOCK
The owner has stopped, retries are disabled, the exact backend is proven,
the state serial is stable and the Terraform lock ID is preserved.
RECOVER STATE
A writer stopped after changing Azure or state and the new plan is not clean.
Reconstruct configuration, state and live resources before rerun or import.
STOP AND ESCALATE
Two writers may have overlapped, lineage or serial is unexpected,
ownership is ambiguous, or recovery proposes unrelated destruction. If two writers actually overlapped, restoring an older blob version is not an automatic rollback. It can erase bindings written by a valid operation. Preserve versions, compare serial and lineage, and use a dedicated state-recovery review.
Prevent the next collision
Locking is the last line of defense, not the pipeline scheduler. Serialize operations by the state identity: storage account, container, key and workspace. Plans, drift checks and applies that can acquire the same state lock belong to the same concurrency group.
Keep production applies single-writer, make queued jobs visible, cancel superseded plans, and prevent blind automatic retries after interruption. Include backend identity in job logs without printing credentials. Use separate state keys where stacks have genuinely separate ownership and lifecycle, not merely to avoid a queue.
After recovery, confirm that one reviewed apply completes, state serial advances as expected, a subsequent plan contains no unexplained action, and the affected Azure service passes its operational probe. That is the validation point. A released lock alone proves nothing about infrastructure health.
Conclusion
A Terraform state lock incident is a concurrency incident before it is a command-line problem. The safe sequence is to pin the exact Azure Blob state, freeze contenders, prove writer liveness, preserve state metadata, remove only a stale Terraform lock, then re-enter through one clean plan.
Wait when a writer is still active. Force-unlock when its absence and the stable state are proven. Move to state recovery when Azure, configuration and state no longer converge. The useful outcome is not an unlocked backend; it is a restored single-writer system with a reviewed plan and an explicit return path.