Infrastructure

Azure Monitor: contain log alert fan-out before changing the KQL rule

A production runbook for measuring split-by dimension cardinality, containing notification noise, canarying stable grouping and rolling back an Azure Monitor log alert without losing coverage.

25 Sept 2026 azureazure-monitorlog-alertsscheduled-query-ruleskqldimensionsobservabilityincidentcanaryrunbookrollbackproduction

One Azure Monitor log alert suddenly opens hundreds of alert instances. The KQL query still returns the expected errors, but the action group pages once per resource, error code and request identifier. Muting the action group quiets the phone; raising a limit postpones the failure. Neither action explains why one operational signal became hundreds of independently evaluated time series.

The running case is a scheduled query rule for API failures. It splits results by _ResourceId, ErrorCode and OperationId. The first two columns describe a useful failure cohort. OperationId is almost unique per request, so every evaluation creates new combinations that do not share lifecycle or notification suppression. This runbook measures the fan-out, contains the incident, redesigns the grouping and ends with a promote, hold or rollback decision.

Freeze the rule and the incident window

Preserve the deployed rule before editing it in the portal. Record its resource ID, enabled state, scopes, query, evaluation frequency, window, threshold, failing periods, dimensions, actions and whether alerts resolve automatically. Also freeze a UTC window containing normal traffic and the alert storm.

bash 01-export-scheduled-query-rule.sh
RG="rg-observability-prod"
RULE="api-errors-prod"

az monitor scheduled-query show --resource-group "$RG" --name "$RULE" --output json > scheduled-query-before.json

az monitor scheduled-query show --resource-group "$RG" --name "$RULE" --query "{enabled:enabled,scopes:scopes,evaluationFrequency:evaluationFrequency,windowSize:windowSize,autoMitigate:autoMitigate,criteria:criteria,actions:actions}" --output yaml

Do not start by changing the query and dimensions together. A query change can alter the population while a dimension change alters the number of independent alert instances. Keep those hypotheses separate.

Reconstruct the fan-out contract

Split-by dimensions are not display labels. Azure Monitor groups query results by every selected combination and evaluates each group separately. Up to six dimensions can be configured, but six low-cardinality columns and six request-level identifiers have radically different operational effects.

Write the intended contract before reading the incident data.

yaml alert-cardinality-contract.yml
signal: failed API requests
decision: page when one service and error family sustain failures
stable_dimensions:
- _ResourceId
- ErrorCode
forbidden_dimensions:
- OperationId
- RequestId
- TimeGenerated
expected_series:
normal: 8-20
incident: less-than-60
notification_budget:
paging_instances_per-evaluation: 5
rollback:
redeploy scheduled-query-before.json
keep notification suppression until the previous rule is healthy

The exact budget belongs to the service, not to a generic platform threshold. Its purpose is to make an unexpected multiplication visible before the action group becomes the first detector.

Measure combinations with the production window

Replay the alert query over the frozen window, but return the grouping columns rather than only the final threshold. Count distinct values independently and then count their combinations.

kusto 02-measure-dimension-cardinality.kql
let Start = datetime(2026-09-25T07:30:00Z);
let End = datetime(2026-09-25T08:00:00Z);
AppRequests
| where TimeGenerated between (Start .. End)
| where Success == false
| extend ErrorCode = tostring(ResultCode)
| summarize
  Rows=count(),
  Resources=dcount(_ResourceId),
  ErrorCodes=dcount(ErrorCode),
  Operations=dcount(OperationId),
  Combinations=dcount(strcat(_ResourceId, "|", ErrorCode, "|", OperationId))

Then inspect which column drives growth and whether it is stable across evaluations.

kusto 03-profile-alert-series.kql
let Start = datetime(2026-09-25T07:30:00Z);
let End = datetime(2026-09-25T08:00:00Z);
AppRequests
| where TimeGenerated between (Start .. End)
| where Success == false
| extend ErrorCode = tostring(ResultCode)
| summarize Hits=count(), FirstSeen=min(TimeGenerated), LastSeen=max(TimeGenerated)
  by _ResourceId, ErrorCode, OperationId
| order by Hits desc

If almost every row has a different OperationId, it is correlation evidence, not a grouping dimension. Keep it in logs and alert payload context where useful, but do not make it the identity of the alert instance.

Separate evaluation, lifecycle and notification

Three mechanisms can look like the same storm.

First, the query can produce more real cohorts because a deployment created failures across many resources. Second, volatile dimensions can turn each row into a new series even when the underlying incident is one failure mode. Third, one stable alert instance can notify repeatedly because resolution, action processing or receiver behavior is wrong.

Compare the number of query rows, unique dimension combinations, active alert instances and delivered notifications over the same timestamps. If combinations rise with rows, fix grouping. If combinations remain stable but notifications repeat, inspect action groups and alert processing rules. If active instances never resolve, inspect autoMitigate, the recovery condition and whether the query continues to return the series after recovery.

Resource Health for the scheduled query rule adds another boundary. A syntax error, semantic error, oversized response, resource-intensive query, validation error or active-alert limit is a rule-health problem, not an empty signal. Preserve that status before disabling anything.

Contain noise without claiming the signal is fixed

If paging is preventing response, use a time-bounded alert processing rule or detach only the paging receiver through the governed deployment path. Keep ticketing or a non-paging evidence channel when possible. Record owner, expiry and the exact rule scope.

Notification suppression does not reduce dimension cardinality, evaluation load or the number of alert instances. Conversely, disabling the scheduled query rule stops evaluation and creates a coverage gap. Use it only when the rule itself threatens the alerting plane and a replacement signal is active.

text containment-decision.txt
Suppress notifications temporarily
signal remains queryable
incident team has another live channel
suppression has owner and expiry

Disable the rule temporarily
rule health or active-instance growth threatens alerting
replacement coverage is already active
exported configuration is retained

Do neither
notifications are actionable and within budget
fan-out represents distinct affected resources

Design grouping around an operational decision

A useful dimension changes ownership, impact or response. _ResourceId, region, deployment ring or a bounded error family can meet that test. A GUID, timestamp, request ID or raw exception message usually does not.

For the running case, aggregate requests into service and error-family cohorts. Keep a representative correlation ID outside the grouping, or retrieve examples from logs after the alert fires.

kusto 04-stable-alert-query.kql
AppRequests
| where TimeGenerated > ago(15m)
| where Success == false
| extend ErrorCode = tostring(ResultCode)
| extend ErrorFamily = case(
  ErrorCode startswith "5", "server",
  ErrorCode == "429", "throttle",
  ErrorCode startswith "4", "client",
  "other")
| summarize FailedRequests=count() by _ResourceId, ErrorFamily

The alert rule should split only on _ResourceId and ErrorFamily. The time clause stays directly in the alert query. Avoid volatile values and unsupported alert-query operators, and ensure every dimension is returned as a string or numeric column.

Canary the rule as a second signal

Do not replace a noisy production rule with an unobserved candidate. Deploy a canary with one non-paging action, a bounded scope or one known resource, and the same evaluation frequency and window. Give it a distinct name and ownership marker.

Run three tests:

  1. A known failure inside the selected cohort must open one instance with the correct resource and error family.
  2. Several requests with different operation IDs but the same failure cohort must remain one alert instance.
  3. Recovery must resolve the instance according to the intended lifecycle.

Observe several complete evaluation and recovery cycles. Compare query totals, dimension combinations, fired instances, resolved instances and notifications between old and candidate rules. Do not judge the canary only by whether a message arrived.

Promote, hold or roll back

Promote the candidate when it detects every required cohort, stays within the cardinality budget, resolves predictably and preserves resource ownership in the payload. Remove temporary notification suppression only after the promoted rule completes a clean evaluation and a controlled fire-and-resolve cycle.

Hold when the query is correct but the dimension budget or recovery behavior is still uncertain. Keep the canary non-paging and leave the previous rule as the authoritative signal.

Roll back when stable cohorts are merged, affected resources disappear from alert context, evaluation health degrades or the candidate misses a positive test. Disable the candidate, redeploy the exported rule or the last approved IaC revision, then verify enabled state, actions, dimensions and Resource Health. A rollback is complete only when coverage and notification routing are both restored.

Conclusion

A log alert storm is not automatically a bad KQL query. The query can be semantically correct while one volatile split-by dimension creates a new operational object for every request. Measure rows, cardinality, instances and notifications separately before changing thresholds or raising limits.

The final decision is observable: retain genuinely distinct resource alerts, replace volatile grouping with stable cohorts, hold the candidate for more evidence, or restore the previous rule. The goal is not fewer alerts at any cost; it is one alert instance per decision the on-call team can actually make.