Automation

Microsoft Sentinel: validate a remediation playbook before production action

A production runbook for bounding a Microsoft Sentinel playbook across trigger, entities, identity, dry run, approval, traces, canary and rollback before any remediation action.

02 Sept 2026 azuremicrosoft-sentinellogic-appssecurityautomationmanaged-identityincident-responseobservabilityguardrailsrunbookrollbackproduction

A Microsoft Sentinel incident contains an identity, a host and an IP address considered suspicious. The playbook already enriches the incident and notifies the team. The next step looks straightforward: disable the account, isolate the host or block the IP automatically. It is also the point where a correlation error becomes a production action.

This use case covers an Azure Logic Apps playbook started by a Sentinel automation rule. It receives the incident, qualifies its entities, collects evidence and hands a remediation to an authorized tool. The goal is not to automate every security response. It is to decide whether this specific playbook should remain enrichment-only, prepare an action, request approval, or execute a bounded remediation with verifiable rollback.

Freeze the contract before the workflow

A remediation playbook must not infer its scope from an incident title or a free-text field. Define the trigger, accepted entities, minimum evidence, allowed action and return path.

yaml sentinel-remediation-contract.yml
playbook: sentinel-contain-compromised-entity-prod
trigger:
type: microsoft_sentinel_incident
events: [created, updated]
automation_rule: ar-high-confidence-containment

accepted_entities:
- account_with_tenant_and_object_id
- host_with_stable_device_or_resource_id
- ip_with_observed_direction_and_window

required_evidence:
- incident_id
- analytics_rule_id
- alert_ids
- entity_identifiers
- first_seen_utc
- last_seen_utc
- source_product
- confidence_reason

execution_modes:
observe: enrich_and_comment
propose: build_action_without_state_change
enforce: execute_after_policy_and_approval

reject_when:
- entity_identifier_is_ambiguous
- incident_is_closed_or_test_data
- target_scope_is_not_production_approved
- same_action_is_already_running
- rollback_or_expiry_is_missing

This contract separates intent from connector mechanics. An incident can contain several accounts with the same display name, an IP shared behind NAT, or a reused hostname. A sensitive action must start from a stable identifier and context, not the first string that resembles an entity.

Separate triggering from decision

An automation rule can invoke a playbook when an incident is created or updated. An update is not always new evidence: an owner change, a comment, or the playbook’s own response can modify the incident. Without a guard, the workflow can call itself again or apply the same action twice.

Build an idempotency key from incident, action type, stable target and contract version. Store it before any write. The playbook should return already processed when a successful or still-running execution owns the same key.

text remediation-idempotency.txt
Idempotency key
incident_id + action_type + target_id + contract_version

Before state change
Reject missing or duplicate target identifiers
Read current target state
Check active execution with the same key
Record requested action and expiry

After state change
Record external operation ID
Record resulting target state
Add a bounded comment to the Sentinel incident
Never retrigger enforcement from that comment alone

The trigger carries the incident. It does not make the decision. Automation-rule conditions reduce noise, while the playbook validates state, entities and policy again immediately before action.

Prove every identity involved

Three permission layers are often confused: the operator configuring the rule, the Microsoft Sentinel service allowed to start the playbook, and the workflow identity calling Sentinel and remediation systems. Broadening them together makes the audit trail unreadable.

Microsoft Sentinel must be permitted to execute playbooks in the relevant resource group. The workflow should use a managed identity or an explicitly inventoried connection, with access limited to its actions. A playbook that adds an incident comment does not require the same role as a connector that disables an identity or isolates an endpoint.

bash inventory-playbook-identities.sh
RG="rg-secops-automation-prod"
LOGIC_APP="la-sentinel-containment-prod"

# Adapt the resource type for the deployed Consumption or Standard workflow.
az resource show \
--resource-group "$RG" \
--name "$LOGIC_APP" \
--resource-type Microsoft.Logic/workflows \
--query "{id:id, identity:identity, state:properties.state}" \
--output json

PRINCIPAL_ID="<managed-identity-principal-id>"
az role assignment list \
--assignee "$PRINCIPAL_ID" \
--all \
--query "[].{role:roleDefinitionName, scope:scope}" \
--output table

The expected output is a readable matrix of identity, role, scope, connector and action. Reject promotion if a personal secret, shared connection or subscription-level role prevents you from attributing the action precisely.

Test observation, then proposal

The first test must not change an account, host or network rule. Run the playbook manually on a realistic test incident and enable enrichment only: read alerts and entities, resolve identifiers, add a comment, record the run and calculate the decision.

The propose mode then builds the exact payload that would be sent to the remediation tool. It validates schema, target, duration, reason, approver and rollback, but replaces the write call with a trace.

json proposed-remediation.json
{
"incidentId": "inc-20260902-042",
"action": "temporary_containment",
"target": {
  "type": "host",
  "id": "stable-resource-or-device-id"
},
"reason": "high-confidence incident with two correlated alerts",
"requestedAtUtc": "2026-09-02T06:20:00Z",
"expiresAtUtc": "2026-09-02T08:20:00Z",
"approval": {
  "required": true,
  "status": "pending"
},
"rollback": {
  "action": "remove_temporary_containment",
  "owner": "secops-on-call"
},
"mode": "propose"
}

Test at least one valid entity, one ambiguous entity, a closed incident, a repeated update, missing permission, an unavailable dependency and a partial response. The error branch must end as failed or held, never as a silent success.

Read incidents and runs over the same window

Validation must connect the Sentinel signal, Logic Apps run history and target tool. In Log Analytics, start with the incident and its alerts; adapt tables and fields to the data connectors enabled in the workspace.

kusto sentinel-remediation-evidence.kql
let Window = 2h;
let TargetIncidentNumber = 42;
let Incident =
SecurityIncident
| where TimeGenerated > ago(Window)
| where IncidentNumber == TargetIncidentNumber
| summarize arg_max(TimeGenerated, *) by IncidentNumber
| project IncidentNumber, Title, Severity, Status, ProviderIncidentId,
        AlertIds, Owner, Labels, LastModifiedTime;
let Alerts =
SecurityAlert
| where TimeGenerated > ago(Window)
| project AlertTime=TimeGenerated, SystemAlertId, AlertName,
        AlertSeverity, ProductName, Entities, ExtendedProperties;
Incident
| mv-expand AlertId = AlertIds
| extend AlertId = tostring(AlertId)
| join kind=leftouter Alerts on $left.AlertId == $right.SystemAlertId
| project IncidentNumber, ProviderIncidentId, Title, Severity, Status,
        AlertTime, AlertName, AlertSeverity, ProductName, Entities
| order by AlertTime asc

The result should support a single timeline: incident received, evidence read, decision calculated, approval, external call, resulting state and any rollback.

Logic Apps diagnostics should retain run ID, executed branch, response code and duration without logging tokens, secrets or unnecessary sensitive content. The target tool must return its own operation identifier. A Succeeded Logic Apps run does not prove that the target reached the expected state.

Canary the production authorization

Do not jump from a test incident to every critical alert. Start with an identifiable canary: one analytics rule, one entity type, one environment and an on-call window. Keep human approval for every state-changing action.

yaml sentinel-playbook-promotion-gate.yml
promotion_gate:
automation_rule_scope:
  analytics_rules: [rule-approved-high-confidence]
  severities: [High]
  entity_types: [Host]
  environments: [production-canary]

require:
  - stable target identifier
  - two independent evidence signals
  - no duplicate execution
  - managed identity at bounded scope
  - approved action payload
  - target operation ID returned
  - post-action state verified
  - expiry or rollback scheduled

stop_when:
  - entity resolution is ambiguous
  - false-positive rate exceeds the agreed threshold
  - target API returns partial or unknown state
  - incident update causes a duplicate run
  - audit trail cannot join incident, run and target operation

Measure selection precision and execution reliability separately. A playbook can call the wrong host perfectly; conversely, a sound decision can fail because of permission. Those risks do not have the same fix.

Decide, validate or roll back

Promotion is acceptable when the playbook rejects ambiguous entities, remains idempotent, uses a bounded identity, enforces the intended approval and proves final state in the target system. Keep a version of both workflow and automation rule with each decision.

If the canary triggers an unjustified action or loses traceability, disable the execution step or automation rule first, then reverse temporary remediations from the recorded operation IDs. Do not delete the runs or comments needed to reconstruct the incident. Return to observe mode while the defect is corrected.

Conclusion

A Microsoft Sentinel playbook is production-ready only when its trigger, decision and action can be explained independently. Entity contracts, idempotency, identities, proposal mode and target-side evidence matter more than a workflow marked Succeeded.

The final decision is explicit: remain in enrichment, prepare an approved action, open a bounded canary, or return to observation. Until target, effective permission and rollback are proven, the playbook must not modify production.