Cloud

Azure Resource Locks: diagnose a blocked deployment before removing the lock

A production runbook for separating inherited Azure management locks, RBAC, Policy and deny assignments, then validating a bounded unlock, deployment and rollback.

30 Sept 2026 azureresource-locksarmbicepterraformdevopsrbacazure-policyactivity-logautomationrunbookrollbackproduction

A production deployment updates several resources, then fails while replacing a diagnostic setting. The pipeline reports ScopeLocked. An operator finds a CanNotDelete lock on the resource group and proposes removing it, rerunning the whole deployment, then recreating the lock.

That sequence may restore the release. It may also expose every resource in the group to deletion, replay operations that already succeeded, or recreate a lock with a different scope and no evidence of what happened while it was absent. The running case is a Bicep or Terraform deployment that must replace one extension resource inside a protected resource group. The runbook must end with a controlled decision: avoid the delete, perform a bounded unlock, stop for another control-plane owner, or roll back a partial deployment.

Freeze the failed operation, not just the job

Preserve the pipeline run, commit, deployment name, UTC window, caller object ID, correlation ID and complete target resource ID. A red job does not tell you which control-plane action was refused or which earlier actions succeeded.

bash 01-capture-deployment-failure.sh
RG="rg-platform-prod"
DEPLOYMENT="platform-20260930-01"
START_UTC="2026-09-30T05:45:00Z"
END_UTC="2026-09-30T06:15:00Z"

az deployment operation group list --resource-group "$RG" --name "$DEPLOYMENT" --query "[].{state:properties.provisioningState,resource:properties.targetResource.resourceName,type:properties.targetResource.resourceType,status:properties.statusMessage}" --output json > deployment-operations.json

az monitor activity-log list --resource-group "$RG" --start-time "$START_UTC" --end-time "$END_UTC" --query "[].{time:eventTimestamp,status:status.value,operation:operationName.value,resource:resourceId,caller:caller,correlationId:correlationId}" --output json > activity-log.json

For Terraform, retain the plan, state serial, provider lock file and complete apply output. Record successful writes separately from the failed action. Do not rerun while the actual state and the deployment engine’s state may disagree.

The useful incident statement is precise: “the pipeline principal attempted DELETE on this diagnostic setting at 06:02 UTC; Azure returned ScopeLocked under this correlation ID after two unrelated updates succeeded.”

Prove which control blocked the request

A failed Azure write can come from several controls. They do not share the same repair:

  • ScopeLocked or an explicit management-lock message points to Microsoft.Authorization/locks.
  • AuthorizationFailed points first to the caller, action and RBAC scope.
  • RequestDisallowedByPolicy must be traced to a Policy assignment and effect.
  • a deny assignment can override an otherwise valid role, including one created by a managed application or deployment stack.

Inspect the raw error before changing permissions. Adding Owner does not override a management lock, and deleting a lock does not repair an RBAC or Policy denial.

bash 02-freeze-caller-and-controls.sh
CALLER_OBJECT_ID="<pipeline-principal-object-id>"
TARGET_ID="/subscriptions/<sub>/resourceGroups/rg-platform-prod/providers/Microsoft.Storage/storageAccounts/stplatformprod/providers/Microsoft.Insights/diagnosticSettings/send-platform-logs"

az ad sp show --id "$CALLER_OBJECT_ID" --query "{id:id,appId:appId,displayName:displayName}" --output json

az role assignment list --assignee-object-id "$CALLER_OBJECT_ID" --all --include-inherited --query "[].{role:roleDefinitionName,scope:scope,condition:condition}" --output json > caller-rbac.json

az lock list --output json > subscription-locks.json
az lock list --resource-group "rg-platform-prod" --output json > resource-group-locks.json

If the target is an extension resource such as a diagnostic setting, remember that it inherits the lock of the resource to which it is attached. A lock may therefore block deletion even when no lock is visible on the extension resource itself.

Reconstruct the effective lock chain

Azure management locks exist at subscription, resource-group or resource scope. Child resources inherit parent locks, and the most restrictive lock in the chain wins. CanNotDelete permits updates but blocks deletion. ReadOnly blocks control-plane updates as well as deletion.

Read every candidate lock as a tuple:

text effective-lock-record.txt
lock_id
lock_name
level: CanNotDelete | ReadOnly
declared_scope
inherited_by_target: true | false
owner_or_change_record
notes
creation_or_last-change evidence
required operation: PUT | PATCH | POST | DELETE

The last field matters. A deployment that replaces a child resource can contain a delete even when the desired configuration looks like an update. Conversely, a CanNotDelete lock cannot explain a failed in-place PUT. Do not infer the operation from the template diff alone; verify the deployment operation or Activity Log.

Locks apply to Azure control-plane calls, not to service data-plane operations. Removing a resource lock is therefore not a repair for a blob, SQL row or queue message operation. Keep the diagnosis on the plane that actually failed.

Decide whether deletion is really required

Regenerate a plan or what-if against the frozen input and identify the exact destructive action.

bash 03-review-destructive-change.sh
RG="rg-platform-prod"
TEMPLATE="main.bicep"
PARAMETERS="main.prod.bicepparam"

az deployment group what-if --resource-group "$RG" --template-file "$TEMPLATE" --parameters "$PARAMETERS" --result-format FullResourcePayloads --output json > what-if.json

jq -r '
.changes[]
| select(.changeType == "Delete" or .changeType == "Create" or .changeType == "Modify")
| [.changeType, .resourceId]
| @tsv
' what-if.json

Classify the change before touching the lock:

  • If a stable in-place update is supported, change the deployment rather than opening the scope.
  • If replacement comes from a renamed IaC resource, provider regression or ownership drift, repair that contract first.
  • If deletion is intentional, identify every dependency and the smallest scope whose lock must change.
  • If the lock belongs to a managed application or another platform owner, stop. The service lifecycle or owning team must perform the operation.

A targeted Terraform apply is not a generic escape hatch. It can help isolate an exceptional repair, but a complete plan must follow to prove that the rest of the graph is still converged.

Build a bounded unlock record

Before removal, export the lock and make the change window explicit. The record must name the exact lock ID, target operation, approver, operator, start and expiry, concurrent deployments to pause, validation owner and recreation command.

For the running case, the lock is deliberately at resource-group scope. Removing it affects more than the diagnostic setting, so freeze other deployment identities for the window and monitor all writes to the group.

bash 04-export-resource-group-lock.sh
RG="rg-platform-prod"
LOCK_NAME="protect-platform-prod"
LOCK_ID="$(az lock show --name "$LOCK_NAME" --resource-group "$RG" --query id -o tsv)"

az lock show --ids "$LOCK_ID" --output json > lock-before.json

jq -e '
.level == "CanNotDelete"
and .name == "protect-platform-prod"
' lock-before.json >/dev/null

az lock delete --ids "$LOCK_ID"

Do not delete every lock returned by a list command. Do not widen the pipeline identity so it can manage locks permanently. The unlock identity and deployment identity should remain separable, with their actions visible in the Activity Log.

An exit trap can reduce human error in an automation wrapper, but it is not the only recovery mechanism: a killed runner or lost credential can still leave the scope unlocked. Use an external timer or monitored change record that alerts before the approved window expires.

Execute one controlled change

Run only the approved deployment artifact and keep the resource list stable. Reject the run if a fresh plan introduces another delete, a different caller is active, or the lock window is nearly exhausted.

Observe control-plane events while the change runs. For the target resource, require the expected operation, caller, correlation ID and final provisioning state. For the resource group, flag every write that does not belong to the approved correlation ID.

Then validate the service outcome, not merely Succeeded: diagnostic data resumes in the expected destination, its schema and latency remain acceptable, and no unrelated setting or role assignment changed.

Recreate the lock and prove convergence

Reapply the lock immediately after the approved operation, before cleanup or optional follow-up work.

bash 05-recreate-and-verify-lock.sh
RG="rg-platform-prod"
LOCK_NAME="protect-platform-prod"

az lock create --name "$LOCK_NAME" --resource-group "$RG" --lock-type CanNotDelete --notes "Production deletion guard; removal requires approved bounded change"

az lock show --name "$LOCK_NAME" --resource-group "$RG" --query "{id:id,level:level,notes:notes}" --output json > lock-after.json

az deployment group what-if --resource-group "$RG" --template-file main.bicep --parameters main.prod.bicepparam --output json > what-if-after.json

Compare lock-before.json and lock-after.json for scope, level, name and operational notes. A lock recreated one level lower is not equivalent. Run the full IaC plan or what-if again and explain every remaining change.

Do not test protection by deleting a production resource. Validate denial behavior on a disposable canary at the same lock pattern before adopting the procedure, and in production prove that the lock exists, the IaC state converges, the application is healthy and unauthorized lock removal is impossible for the deployment principal.

Decide, validate or roll back

Keep the lock and redesign the deployment when the delete is accidental, an in-place update exists, or the replacement scope is too broad.

Approve a bounded unlock only when the exact delete is intentional, the effective lock and owner are proven, concurrent writers are paused, the prior state is recoverable, monitoring is active and lock recreation has a tested owner.

Stop and escalate when the lock is inherited from a subscription, owned by a managed application, paired with an unexplained deny assignment, or impossible to recreate before the window expires.

If the deployment has partially applied, recreate the lock first unless doing so prevents the documented compensation. Restore the last known configuration, reconcile the IaC state, validate service health, then produce a new full plan. Never use an unlocked scope as permission to rerun a partially understood deployment.

Conclusion

A management lock is a control-plane safety boundary, not an inconvenient checkbox. The incident is resolved only when the blocked operation is identified, the effective lock chain is proven, the change is reduced to an intentional scope, and protection is restored with evidence.

That leaves a defensible production decision: avoid the delete, open one monitored window, hand the operation back to the owning service, or roll back the partial release without weakening the platform permanently.