Infrastructure
Azure Update Manager: diagnose a maintenance window before forcing patching
A production runbook for qualifying an incomplete Azure Update Manager campaign across dynamic scope, orchestration, window budget, installation, reboot and validation before any out-of-window retry.
The maintenance window has ended, yet part of the fleet is still noncompliant. Some VMs show installed updates, others have no recent history, and a few servers are waiting for a restart. The tempting response is an immediate one-time installation across the fleet. That retry can patch machines that were never in the approved scope, run outside the business window, or hide a targeting defect.
This runbook covers Azure VMs and Azure Arc-enabled servers attached to an Azure Update Manager Maintenance Configuration, either directly or through dynamic scope. Its purpose is to decide whether to resume on a bounded set, wait for the next window, repair the configuration, or stop patching and use an application recovery path.
Freeze the maintenance contract
Preserve the configuration that was expected to run, not only the state visible after the incident. A patching window joins a schedule, time zone, duration, update classifications, reboot policy and machine population.
maintenance_configuration: mc-prod-linux-weekly
window:
start_utc: 2026-09-01T01:00:00Z
duration: PT2H
expected_end_utc: 2026-09-01T03:00:00Z
selection:
assignment: dynamic_scope
filters:
resource_groups: [rg-app-prod]
os_types: [Linux]
tags:
patch_ring: ring-2
environment: production
updates:
classifications: [Critical, Security]
reboot_setting: IfRequired
evidence:
- maintenance configuration version
- assignment and dynamic-scope filters
- machine inventory at window start
- pre-window update assessment
- installation and maintenance run IDs
- application health before and after This record prevents the investigation from comparing results with a configuration edited after the fact. It also makes the real execution budget explicit. Update Manager reserves part of the window for completion and reboot work, so a campaign can stop safely before every selected update is installed.
Rebuild the population that actually matched
A dynamic-scope preview is not proof of the runtime population. Membership is evaluated again: a machine created, moved or retagged between preparation and execution can enter or leave the scope.
Build three sets:
- machines approved in the change inventory;
- machines matching the filters during diagnosis;
- machines with a maintenance or installation run in the incident window.
Explain every difference before retrying. Check tag spelling and case, resource group, operating system, Maintenance Configuration assignment and the level where the dynamic scope was created. A resource-group-level assignment does not select machines elsewhere, even in the same subscription.
For Azure VMs, also prove that patch orchestration supports Customer Managed Schedules. A machine visible in inventory but not configured for the schedule is not the same failure as a machine whose patch installation started and failed.
Separate no run, exhausted window and patch failure
Do not put every noncompliant machine into a single “patching failed” bucket. Each family has a different recovery path.
No maintenance run
Check assignment, dynamic scope, schedule, time zone and orchestration
Run created, no installation
Check assessment, classifications, exclusions and applicable updates
Partial installation without package error
Check remaining window time and the reboot reserve
One or more updates failed
Read per-update results and package-manager logs
Installation succeeded, machine still noncompliant
Run a fresh assessment and check replacement or supersedence
Restart required or application health degraded
Do not force another batch; qualify service recovery first A global Succeeded status does not prove that every expected machine was targeted. Conversely, an incomplete run can be a maintenance-window safeguard rather than an update-engine fault.
Read execution evidence in Azure Resource Graph
Azure Update Manager publishes assessment and installation results to Azure Resource Graph. Reconstruct the campaign with run and machine identifiers instead of relying on a compliance counter.
let WindowStart = datetime(2026-09-01T00:30:00Z);
let WindowEnd = datetime(2026-09-01T03:30:00Z);
patchinstallationresources
| where type !has "softwarepatches"
| extend p = parse_json(properties)
| extend LastModified = todatetime(p.lastModifiedDateTime),
Machine = tostring(split(id, "/", 8)),
ResourceGroup = tostring(split(id, "/", 4)),
Status = tostring(p.status),
Installed = toint(p.installedPatchCount),
Failed = toint(p.failedPatchCount),
Pending = toint(p.pendingPatchCount),
Excluded = toint(p.excludedPatchCount),
RebootStatus = tostring(p.rebootStatus)
| where LastModified between (WindowStart .. WindowEnd)
| project LastModified, name, Machine, ResourceGroup,
Status, Installed, Failed, Pending, Excluded, RebootStatus
| order by LastModified asc This query returns only machines that produced an installation result. Compare it with the expected inventory. For an absent machine, also query microsoft.maintenance/applyupdates records in maintenanceresources to determine whether the Maintenance Configuration started work without producing a usable installation result.
Retain per-update detail for failed machines. Package name, classification and installationState help separate repository access, dependency, conflict, exclusion and pending-reboot failures.
Validate the host and the service before resuming
The Azure control-plane result is only part of the evidence. On one canary from each failure family, check clock, disk space, package manager, failed services and reboot requirement. For Azure Arc, include connected-machine and agent state.
Keep application validation independent from patch status:
- health probe and error rate;
- available capacity and instance distribution;
- access to dependencies;
- kernel or component version actually loaded after restart;
- ability to drain one instance before intervention.
A compliant machine that never returned to the load balancer is not a successful campaign. A machine with a pending update may remain healthy until a better prepared window.
Choose a bounded recovery
Do not rerun the complete Maintenance Configuration to repair three hosts. Create an explicit set from the evidence and make a decision for each failure family.
decisions:
resume_bounded:
when:
- affected machines are explicitly listed
- application redundancy is proven
- a new approved window exists
- reboot and stop conditions are defined
wait_next_window:
when:
- patches remain pending without emergency exposure
- current service health is stable
- dynamic scope and orchestration are corrected
stop_and_repair:
when:
- target membership is unexplained
- package or Arc connectivity failures repeat
- maintenance evidence is incomplete
rollback_service:
when:
- post-patch health degrades
- safe OS patch removal is not proven
- traffic can return to a known-good instance or image
stop_conditions:
- error rate exceeds the approved threshold
- healthy capacity drops below quorum
- an unexpected machine enters the batch
- reboot exceeds the per-machine budget
- run IDs cannot be correlated OS patch rollback is not a universal primitive. Some packages can be removed and others cannot; a previous kernel might remain available, but that is not a recovery plan to invent during an incident. The more reliable path is often at the service layer: drain the instance, restore a known image, recover capacity or shift traffic while the host is repaired.
Repair the next window without deleting evidence
After stabilization, fix the cause at the right layer: dynamic-scope filter, assignment, orchestration, classifications, window duration, reboot policy or operating-system prerequisite. Run a fresh assessment first, then use a canary set that can complete installation, reboot and validation within budget.
Keep run IDs, Resource Graph results, configuration versions and application checks with the change record. Final compliance must converge with the expected population and service health, not just a percentage in the portal.
Conclusion
An incomplete Azure Update Manager campaign is not evidence that the entire fleet should be patched again. Rebuild the runtime scope, separate no run, exhausted window, update failure and pending restart, then prove application capacity.
The production decision becomes explicit: resume an identified set, wait for the next window, repair the schedule or return to known-good capacity. If target membership or recovery remains ambiguous, do not force patching outside the approved window.