Infrastructure

Azure Monitor: diagnose suppressed notifications before changing the alert

A production runbook to prove that an alert processing rule suppresses notifications, bound its scope, validate recovery, and keep a rollback path.

29 Aug 2026 azureazure-monitoralertsalert-processing-rulesaction-groupsobservabilityincidentautomationsecurityrunbookrollbackproduction

An availability alert appears as fired in Azure Monitor, but on-call receives no phone call, email, or webhook. Teams often start by changing the alert rule or rebuilding its action group. Yet the signal was evaluated successfully: an alert processing rule may have removed every action group while the alert fired.

The production case is a rule created for planned maintenance that remained active with an overly broad scope or a misread schedule. This runbook separates signal generation, fired-alert processing, and notification delivery. It ends with a bounded decision: keep the suppression, narrow its scope, disable the processing rule, repair the action group, or roll back the latest observability change.

Freeze one fired alert with no notification

Start with one exact alert instance. Preserve its identifier, UTC timestamp, target resource, originating alert rule, severity, and expected action groups. Do not disable several components together; that would make recovery impossible to attribute.

yaml alert-notification-incident.yml
incident:
started_at_utc: 2026-08-29T14:10:00Z
fired_alert_id: <alert-instance-id>
alert_rule_id: <alert-rule-resource-id>
target_resource_id: <monitored-resource-id>
severity: Sev1
expected_action_group_ids:
  - <on-call-action-group-id>

observed:
alert_visible_in_azure_monitor: true
notification_received: false
action_group_test_status: <not-run>

last_changes:
- planned maintenance window
- alert processing rule deployment
- scope or filter update
- time-zone or recurrence update
- action group receiver update

stop_conditions:
- no alert instance can be identified
- the monitored signal did not breach
- suppression scope includes unrelated production resources
- no alternate paging path protects the service

An alert processing rule is not the alert rule. The former changes the actions applied when an alert fires; the latter evaluates a signal and creates the alert. Suppression removes all action groups from the affected fired alert, while the alert remains visible in the portal, API, and Azure Resource Graph. That distinction is the first diagnostic clue.

Prove the three layers independently

The investigation must answer three questions in order:

  • did the signal cross its condition during the expected window;
  • was an alert instance created for the right target and severity;
  • which processing rules could modify its actions at that time.

If no alert exists, stay with the query or metric, evaluation frequency, window, dimensions, or alert-rule state. If the alert exists, do not change its threshold to repair delivery. Rebuild the chain from the fired alert through processing rules to the action group.

bash 01-alert-processing-inventory.sh
SUBSCRIPTION_ID="<subscription-id>"

az account set --subscription "$SUBSCRIPTION_ID"

az monitor alert-processing-rule list \
--query "[].{name:name, resourceGroup:resourceGroup, enabled:properties.enabled, scopes:properties.scopes, actions:properties.actions, conditions:properties.conditions, schedule:properties.schedule}" \
--output json

az resource list \
--resource-type "Microsoft.AlertsManagement/actionRules" \
--query "[].{name:name, resourceGroup:resourceGroup, id:id, location:location}" \
--output table

The resource type keeps its historical name, Microsoft.AlertsManagement/actionRules. Inventorying by type therefore finds rules deployed through IaC even when the team calls them alert processing rules. The Azure CLI command comes from the alertsmanagement extension. Pin that extension in automation and preserve the raw output with the incident evidence.

Calculate the scope that actually applies

A rule can target one resource, several resources, a resource group, or an entire subscription. It only affects alerts fired on resources inside that scope and in the same subscription. Filters can then narrow the set by alert rule name or ID, severity, monitor service, resource, or alert context.

Build an applicability matrix instead of trusting the rule name.

text alert-processing-applicability.txt
Fired alert
target resource ID: <exact-resource-id>
alert rule ID: <exact-alert-rule-id>
alert rule name: <display-name>
severity: Sev1
monitor service: <service>
fired at UTC: <timestamp>

Candidate processing rule
enabled: true-or-false
action: RemoveAllActionGroups-or-AddActionGroups
scope contains target: true-or-false
filters match alert: true-or-false
schedule active at fired time: true-or-false
deployment propagation complete: true-or-false

Applicable only when
enabled AND scope contains target AND every filter matches
AND schedule is active AND propagation is complete

Compare full resource IDs, not display names. A rule named maintenance-app1 can still contain a shared resource group. An empty filter is not a protective filter. Several values inside one condition may broaden the match; several conditions should be reviewed as a contract, not read as a sentence.

Replay the schedule in its time zone

Maintenance suppressions commonly fail at boundaries: local time stored as UTC, a weekly recurrence that crosses midnight, a one-time rule that was never retired, or a schedule left behind after a daylight-saving transition. Always compare the alert’s UTC time with both the rule definition and its configured time zone.

bash 02-show-processing-rule.sh
RULE_RG="rg-monitoring-prod"
RULE_NAME="suppress-maintenance-app1"

az monitor alert-processing-rule show \
--resource-group "$RULE_RG" \
--name "$RULE_NAME" \
--output json > alert-processing-rule.json

jq '{
id,
enabled: .properties.enabled,
scopes: .properties.scopes,
conditions: .properties.conditions,
schedule: .properties.schedule,
actions: .properties.actions,
description: .properties.description,
tags
}' alert-processing-rule.json

A new or updated processing rule can take up to thirty minutes to affect newly fired alerts. Do not treat an alert raised during that propagation period as final evidence. Record the change time, wait for the documented propagation window, and generate a new bounded test signal.

Separate suppression from an action group failure

After qualifying the processing rule, test the action group independently. A receiver test failure points toward the channel, webhook authentication, address, phone number, rate limit, or downstream integration. A successful test does not prove that the action group was attached to the fired alert.

The evidence sequence is:

  1. the production alert exists;
  2. an applicable processing rule uses RemoveAllActionGroups;
  3. the action group succeeds in an independent test;
  4. a new controlled alert notifies after the suppression is narrowed or disabled.

When multiple processing rules apply, suppression takes priority over adding action groups. A second rule that adds an action group therefore cannot override an existing suppression. Repair the rule that removes the actions.

Contain the issue without reopening every notification

When the scope is too broad, do not delete the rule immediately if it still protects active maintenance. Choose the narrowest mitigation:

  • disable the rule when its window is over and its owner confirms retirement;
  • replace subscription scope with the resources actually under maintenance;
  • add a stable filter tied to the intended alert rule or target resource;
  • fix the time zone or recurrence with an explicit window;
  • keep an emergency paging route outside the suppression boundary until validation completes.
bash 03-contain-processing-rule.sh
RULE_RG="rg-monitoring-prod"
RULE_NAME="suppress-maintenance-app1"

# Reversible mitigation when suppression must stop.
az monitor alert-processing-rule update \
--resource-group "$RULE_RG" \
--name "$RULE_NAME" \
--enabled false

az monitor alert-processing-rule show \
--resource-group "$RULE_RG" \
--name "$RULE_NAME" \
--query "{enabled:properties.enabled, scopes:properties.scopes, actions:properties.actions, schedule:properties.schedule}" \
--output json

The update command is suitable for reversible disablement. Change scope, filters, or schedule through the IaC source that owns the resource and review the diff. An unreconciled portal fix is likely to be overwritten by the next deployment.

Validate with a bounded synthetic alert

Do not close the incident because the configuration looks correct. Use a test alert rule or a controlled condition against a non-critical resource, with a validation action group and a short window. The test must create a new alert instance after propagation; an old fired alert does not replay notifications when suppression ends.

yaml alert-notification-change-gate.yml
candidate_change:
processing_rule: suppress-maintenance-app1
expected_state: disabled-or-narrowed
owner: platform-observability
propagation_wait: up-to-30-minutes

synthetic_validation:
target: <non-critical-test-resource>
alert_rule: <synthetic-alert-rule-id>
expected_action_group: <validation-action-group-id>
correlation_id: <change-or-incident-id>

promote_when:
- a new alert instance is visible
- the expected receiver gets exactly one notification
- unrelated maintenance alerts remain suppressed
- the production paging path passes its independent test

rollback_when:
- notification fan-out reaches an unintended scope
- duplicate notifications appear
- the maintenance perimeter is no longer protected
- rule applicability cannot be reconstructed

rollback:
- restore the previous versioned processing rule
- keep the emergency paging path active
- preserve alert IDs, rule definitions and test timestamps

Observe a representative cycle afterward: fired alerts, delivered notifications, duplicates, expected suppressions, and rule changes. The criterion is not simply less silence. Every alert class must have an explainable expected action.

Decide whether to repair, retain, or roll back

Keep the suppression when its scope, filters, and schedule still match approved maintenance and another path protects incidents outside that maintenance. Narrow it when it covers unintended resources or severities. Disable it when the window is over, the owner is unknown, or it hides a critical alert without a compensating control.

Repair the action group only when the alert is not suppressed and its receiver test fails. Roll back the latest observability change when the previous version had a known boundary and the new definition cannot be qualified quickly.

Conclusion

A fired alert with no notification is not automatically a broken alert rule. Azure Monitor may have produced the signal correctly, then removed every action group through an applicable alert processing rule.

The reliable decision comes from layered evidence: alert instance, scope and filters, schedule, action priority, then an independent action group test. Correct the smallest boundary, wait for propagation, validate with a new synthetic alert, and retain the previous definition as an explicit rollback.