Automation
Azure Policy: bound a remediation task before it rewrites production
A production runbook for validating the target set, identity, roles, canary, logs and rollback of an Azure Policy remediation before scaling it out.
An Azure Policy initiative reports several hundred non-compliant resources. The rule looks straightforward: add missing diagnostic settings with deployIfNotExists, or correct a property with modify. Create remediation task appears to be the obvious next step. In production, however, it starts automation backed by a managed identity that can write to every resource selected by the assignment.
The use case is a platform team applying a new governance rule to existing resources. The objective is not to slow compliance down. It is to turn a broad remediation into an observable, reversible change: pin the effective definition, produce the exact target list, review the identity permissions, test one representative resource, and scale only when both resource state and telemetry prove the expected effect.
Freeze the contract before creating the task
A remediation task does not execute a display name. It executes the effect of a definition, with assignment parameters, at a given scope. For an initiative, policyDefinitionReferenceId selects the individual rule to remediate. Capture those inputs before execution because a definition or assignment changed between review and launch changes the actual contract.
Change: chg-20260814-021
Assignment: enforce-platform-baseline-prod
Scope: /subscriptions/00000000-0000-0000-0000-000000000000
Definition reference: deploy-diagnostics-storage
Effect: deployIfNotExists
Parameters: target workspace, categories, allowed regions
Initial discovery mode: ExistingNonCompliant
Expected state
Add missing diagnostic settings
Do not replace an approved existing configuration
Touch only in-scope Storage resources
Validation
One canary resource first
Logs arrive in the expected workspace
No role or parameter is widened during execution
Rollback
Cancel a task that is still running
Remove only objects created by the faulty remediation
Restore the versioned definition or assignment if it drifted Keep the definition ID and repository version, assignment parameters, exemptions, scope and identity. A NonCompliant state does not yet describe the write that will occur.
Read the effect and the effective roles
The deployIfNotExists and modify effects use the managed identity attached to the assignment to deploy or change existing resources. The definition’s roleDefinitionIds describes required roles, but it does not prove that the identity received those roles at the smallest useful scope. A subscription-level Contributor assignment deserves review even when the policy remediates a single resource type.
ASSIGNMENT_ID="/subscriptions/00000000-0000-0000-0000-000000000000/providers/Microsoft.Authorization/policyAssignments/enforce-platform-baseline-prod"
az policy assignment show --ids "$ASSIGNMENT_ID" --query '{name:name,scope:scope,definition:policyDefinitionId,parameters:parameters,identity:identity,location:location}' --output json
PRINCIPAL_ID="00000000-0000-0000-0000-000000000000"
az role assignment list --assignee "$PRINCIPAL_ID" --all --query '[].{role:roleDefinitionName,scope:scope}' --output table For a custom definition, inspect the deployIfNotExists template or each modify operation, the existence condition and roleDefinitionIds. Editing a definition does not automatically realign permissions on the assignment identity. The review must compare three objects: definition, assignment and effective role assignments.
Materialize the eligible resource set
Aggregate compliance is not enough to approve a write. Query Azure Policy states through Resource Graph to recover resource IDs, regions, types and evaluation timestamps. Remove exemptions, environments outside the change window and resources whose owner has not accepted the operation.
let AssignmentId = tolower("/subscriptions/00000000-0000-0000-0000-000000000000/providers/Microsoft.Authorization/policyAssignments/enforce-platform-baseline-prod");
let DefinitionReferenceId = "deploy-diagnostics-storage";
policyresources
| where type =~ "microsoft.policyinsights/policystates"
| extend assignmentId = tolower(tostring(properties.policyAssignmentId)),
definitionReferenceId = tostring(properties.policyDefinitionReferenceId),
complianceState = tostring(properties.complianceState),
targetId = tolower(tostring(properties.resourceId)),
targetType = tostring(properties.resourceType),
targetLocation = tostring(properties.resourceLocation),
evaluatedAt = todatetime(properties.timestamp)
| where assignmentId == AssignmentId
| where definitionReferenceId == DefinitionReferenceId
| where complianceState == "NonCompliant"
| project targetId, targetType, targetLocation, evaluatedAt
| order by targetLocation asc, targetId asc Export that list into the change record and calculate a digest. A fresh evaluation can change the target set during the window. ExistingNonCompliant uses the known compliance state, whereas ReEvaluateCompliance starts a new scan before remediation. Make that choice explicit, especially after a recent definition update.
Validate the write before broad remediation
For deployIfNotExists, validate the embedded ARM template separately with what-if in a representative environment. For modify, review every operation and alias: it can add or replace a value, or become inapplicable depending on alias modifiability. In both cases, start remediation on one named canary resource, not an arbitrary numeric subset at subscription scope.
ASSIGNMENT_ID="/subscriptions/00000000-0000-0000-0000-000000000000/providers/Microsoft.Authorization/policyAssignments/enforce-platform-baseline-prod"
CANARY_ID="/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-policy-canary/providers/Microsoft.Storage/storageAccounts/stpolicycanary01"
TASK="remediate-diagnostics-canary-20260814"
az policy remediation create --resource "$CANARY_ID" --name "$TASK" --policy-assignment "$ASSIGNMENT_ID" --definition-reference-id deploy-diagnostics-storage --resource-discovery-mode ExistingNonCompliant
az policy remediation show --resource "$CANARY_ID" --name "$TASK" --output json The canary must share the same type, region, locks and network constraints as later targets. An empty resource created only to make the test pass does not validate existing configuration conflicts or effective permissions.
Correlate task, deployments and Activity Log
A global Succeeded state says the task finished. It does not prove the intended operational result. Capture the remediation correlationId, deployment summary and operations performed by the managed identity. For diagnostic settings, confirm that expected categories reach the target workspace. For modify, read the resulting property from the resource.
let Start = datetime(2026-08-14T07:30:00Z);
let End = datetime(2026-08-14T08:00:00Z);
let RemediationIdentity = "00000000-0000-0000-0000-000000000000";
AzureActivity
| where TimeGenerated between (Start .. End)
| where Caller == RemediationIdentity
| project TimeGenerated,
OperationNameValue,
ActivityStatusValue,
ResourceId,
ResourceGroup,
CorrelationId,
Properties
| order by TimeGenerated asc Retain failures as well as successes. Missing permissions, resource locks, unsupported regions or naming conflicts can create a partial outcome. A task can continue while individual deployments fail, so batch shape and failure threshold must be decided before expanding the scope.
Scale through controlled batches
After the canary is validated, use a child scope or coherent resource group, low concurrency and a bounded resource count. PowerShell exposes ResourceCount, ParallelDeploymentCount, FailureThreshold and ResourceDiscoveryMode, making pace and stop conditions visible in the change request.
$params = @{
Name = "remediate-diagnostics-weu-batch01"
PolicyAssignmentId = "/subscriptions/00000000-0000-0000-0000-000000000000/providers/Microsoft.Authorization/policyAssignments/enforce-platform-baseline-prod"
PolicyDefinitionReferenceId = "deploy-diagnostics-storage"
Scope = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-data-prod-weu"
ResourceDiscoveryMode = "ExistingNonCompliant"
ResourceCount = 20
ParallelDeploymentCount = 2
FailureThreshold = 0.10
}
Start-AzPolicyRemediation @params A batch is accepted only when its before-and-after inventory matches, every write is attributable, and service signals remain healthy. Do not increase scope, concurrency and resource count in the same step; doing so removes the variable that explains a failure.
Decide whether to continue, cancel or compensate
Canceling a task prevents additional operations but does not undo changes already applied. Rollback therefore depends on the effect. A modified property must return to its known previous value. A deployed object should be removed only when its origin, lack of consumers and owner are proven. Switching the assignment to audit can prevent additional automatic changes during investigation, but it does not restore state.
Continue
Canary is compliant and operational effect is proven
Targets match the approved inventory
Identity and roles are bounded
No unexplained failure remains
Hold expansion
Partial result or delayed evidence
New targets appeared since the last evaluation
Existing properties were replaced unexpectedly
Cancel the task
Wrong definition reference
Scope or identity is too broad
Writes occurred outside the inventory
Failure rate exceeds the accepted threshold
Compensate
Restore the previous value for modify
Remove only a created deployment with no consumer
Revalidate compliance and service behavior
Preserve correlationId, touched resources and decision Remediation is complete when technical state, service signal and target inventory converge. A higher compliance percentage is not sufficient if logs remain empty, an approved configuration was overwritten, or the identity retains unnecessary privileges.
Conclusion
An Azure Policy remediation task is an identity-backed production operation. Before remediating broadly, pin the definition and assignment, inventory eligible resources, review roles, validate the template or modify operations, then use a canary and observable batches.
The outcome becomes operational: continue when the effect is proven, stop when the inventory drifts, or compensate precisely for writes already applied. Compliance remains automated without becoming an opaque production rewrite.