Automation

Microsoft Sentinel: diagnose a silent scheduled detection before widening its lookback

A production runbook for separating source ingestion, latency, analytics-rule execution, KQL windows and incident creation before changing a Microsoft Sentinel detection.

17 Sept 2026 azuremicrosoft-sentinelsecurityanalytics-ruleskqlingestionobservabilityincident-responseautomationrunbookrollbackproduction

A scheduled Microsoft Sentinel rule has not created an incident since yesterday. The source platform still shows security events, the connector is marked connected, and running a broad KQL query manually returns rows. Extending the rule lookback from fifteen minutes to several hours looks like a quick recovery.

That change can hide the real failure and replay old events into duplicate alerts. The missing incident may come from a collection gap, delayed ingestion, a schema change, a failed rule execution, a query filter, entity mapping, alert grouping or an automation rule that closes incidents. This runbook isolates each boundary before changing production detection logic.

Freeze one expected detection

Choose one event that should have matched and record its source timestamp, stable identifier and expected entity. Also capture the rule resource ID, current query, frequency, lookback, threshold, grouping and incident settings. Without one known event, a manual query returning any row proves very little.

yaml sentinel-detection-contract.yml
workspace: law-sec-prod
rule: detect-admin-group-change
rule_kind: scheduled
source_table: CommonSecurityLog

expected_event:
source_time_utc: 2026-09-17T06:18:00Z
source_event_id: evt-78421
entity: admin@example.com

schedule:
frequency: 5m
lookback: 15m
threshold: greater-than-0

capture_before_change:
- rule-resource-id-and-etag
- deployed-query-and-parameters
- connector-and-rule-health
- alert-grouping-and-incident-settings

rollback_on:
- duplicate-alerts
- query-cost-or-runtime-regression
- loss-of-entity-mapping
- unexplained-volume-increase

Export the deployed rule definition or preserve the reviewed IaC version. Treat the portal view as evidence of current state, not as the only rollback copy.

Prove source, arrival and searchable row separately

First verify that the source emitted the selected event. Then find the same identifier in the target Sentinel table without reusing all filters from the rule. Finally compare source time, TimeGenerated and ingestion_time(). These are three different clocks.

kusto 01-prove-event-arrival-and-latency.kql
let EventId = "evt-78421";
let WindowStart = datetime(2026-09-17T05:45:00Z);
let WindowEnd = datetime(2026-09-17T07:00:00Z);
CommonSecurityLog
| where TimeGenerated between (WindowStart .. WindowEnd)
| where tostring(ExternalID) == EventId
| extend IngestedAt = ingestion_time()
| extend IngestionDelay = IngestedAt - TimeGenerated
| project TimeGenerated, IngestedAt, IngestionDelay,
        ExternalID, DeviceVendor, DeviceProduct,
        SourceUserName, Activity
| order by TimeGenerated asc

Adapt the stable identifier and columns to the actual table. If the row is absent, stop changing the analytics rule: the failure is upstream or at ingestion. If ingestion_time() is null, do not manufacture a latency conclusion; the function depends on ingestion-time policy. If the row arrived after the rule window, quantify the delay distribution instead of reasoning from one event.

Measure the gap before extending any window

Compare event volume and ingestion delay before and after the last known successful detection. Percentiles are more useful than an average because a small delayed tail can contain the event that matters.

kusto 02-measure-ingestion-delay.kql
let Horizon = 24h;
CommonSecurityLog
| where TimeGenerated >= ago(Horizon)
| extend IngestedAt = ingestion_time()
| where isnotnull(IngestedAt)
| extend DelaySeconds = datetime_diff("second", IngestedAt, TimeGenerated)
| summarize Events = count(),
          P50 = percentile(DelaySeconds, 50),
          P95 = percentile(DelaySeconds, 95),
          P99 = percentile(DelaySeconds, 99),
          LastEvent = max(TimeGenerated),
          LastIngestion = max(IngestedAt)
by bin(TimeGenerated, 15m)
| order by TimeGenerated asc

A wider lookback may be justified when observed latency exceeds the current window. It must be paired with a strategy that prevents overlapping runs from alerting twice. For scheduled detections, filtering on recent ingestion time can assign late events to one execution while the TimeGenerated range covers the measured delay. Test that behavior on the exact table because ingestion characteristics differ by source.

Check connector health without treating it as complete evidence

Enable Microsoft Sentinel auditing and health monitoring if it is part of the workspace operating model. Query _SentinelHealth() and the Data collection health monitoring workbook, but remember that connector health coverage is not universal. A green connector state does not prove that a specific table, tenant or device is producing the expected rows.

kusto 03-read-sentinel-health.kql
let Start = ago(24h);
_SentinelHealth()
| where TimeGenerated >= Start
| where SentinelResourceType in ("Data connector", "Analytics rule")
| where SentinelResourceName has_any ("detect-admin-group-change", "CommonSecurityLog")
| project TimeGenerated, SentinelResourceType, SentinelResourceKind,
        SentinelResourceName, Status, Description, ExtendedProperties
| order by TimeGenerated desc

Use the health record as one layer. Correlate it with source backlog, agent or API diagnostics, table freshness and Azure Activity changes. SentinelAudit can show that the rule changed; it cannot prove that the new query matches the event contract.

Prove that the analytics rule ran on the expected window

A correct query can still fail as a scheduled rule. Functions may be missing, a cross-workspace reference may be unavailable, the query may exceed resources, or all retry attempts may fail. Group health events by the query start time to identify a window that never completed successfully.

kusto 04-find-skipped-rule-windows.kql
_SentinelHealth()
| where TimeGenerated >= ago(24h)
| where SentinelResourceType == "Analytics rule"
| where SentinelResourceKind == "Scheduled"
| where SentinelResourceName == "detect-admin-group-change"
| extend QueryStart = todatetime(ExtendedProperties["QueryStartTimeUTC"]),
       ExecutionStart = todatetime(ExtendedProperties["executionStart"])
| summarize Attempts = count(),
          SuccessfulAttempts = countif(Status == "Success"),
          LastSeen = max(TimeGenerated)
by QueryStart, SentinelResourceId
| where SuccessfulAttempts == 0
| order by QueryStart desc

Do not infer failure solely from a single unsuccessful attempt: scheduled rules can retry the same window. Preserve ExtendedProperties for the incident record before editing and saving the rule.

Replay the deployed query in controlled stages

Run the exact deployed query with fixed timestamps that include the selected event. Then remove one layer at a time: threshold aggregation, joins, watchlists, functions and business filters. The first stage that drops the event identifies the boundary to repair.

text sentinel-query-replay-matrix.txt
Stage 1 - source row
Stable event identifier exists in the expected table

Stage 2 - time contract
Fixed TimeGenerated window includes the event

Stage 3 - enrichment
Joins, watchlists and functions keep the event

Stage 4 - detection logic
Business filters and aggregation produce the expected result

Stage 5 - scheduled execution
Rule health shows a successful run for the target window

Stage 6 - alert and incident
Entity mapping, grouping and incident settings produce one canary incident

Stage 7 - downstream automation
Automation rules and playbooks preserve the intended incident state

This order prevents a collection incident from being “fixed” with KQL and prevents an incident-creation setting from being blamed on ingestion.

Validate with a bounded canary

Clone the rule under a canary name or deploy a parameterized test version that targets one controlled event marker. Keep automated remediation disabled. Replay a benign source event or use an approved test source, then capture one chain: source ID, table row, ingestion time, rule execution, alert ID and incident ID.

Validation succeeds only when the event creates exactly one alert and one incident within the measured latency budget, with the expected entities and no unexpected automation. Also run a negative case that must not alert. A positive-only test cannot detect an over-broad repair.

Decide repair, hold or rollback

Repair the connector or source when the row never arrives. Adjust the rule window only when measured ingestion latency exceeds its contract, and add a tested ingestion-time guard to control duplicates. Repair KQL when the fixed replay shows the event disappearing at a known filter or enrichment step. Repair incident settings or downstream automation only after the alert is proven.

Hold the change when health monitoring is absent, the source event cannot be identified, or the deployed rule differs from the reviewed version. Roll back when the canary creates duplicates, expands the result set unexpectedly, loses entity mapping or increases query runtime beyond its operational budget. Restore the captured rule definition, disable the canary and replay both positive and negative cases.

Conclusion

A silent Microsoft Sentinel detection is not automatically a short lookback problem. It is a chain from source emission to ingestion, query execution, alert creation, incident grouping and automation. Prove one expected event at every boundary before editing the rule.

Widen the window only from measured latency and with duplicate control. Otherwise repair the failed layer, validate with one bounded canary and keep the original rule definition ready for rollback.