Infrastructure

Azure VMSS: diagnose an automatic repair loop before changing the action

A production runbook to separate health signal, grace period, service state, VMSS model and local data before changing Replace, Reimage or Restart.

12 Sept 2026 azurevmssvirtual-machine-scale-setsautomatic-repairsapplication-healthobservabilityautomationinfrastructurerunbookvalidationrollbackproduction

An Azure Virtual Machine Scale Set repeatedly replaces an instance reported as unhealthy. The new VM starts, joins the pool, then gets replaced again. Useful capacity shrinks, and the proposed response is to switch the automatic repair action from Replace to Restart, extend the grace period, or disable repairs.

Each option can stop the loop while preserving its cause. The health signal may probe the wrong endpoint, the VMSS model may recreate the failure, or the application may need longer than the configured grace period. This runbook follows a stateless worker pool using the Application Health extension and a Replace policy. The objective is to preserve capacity, identify the repeatable fault, validate a canary, then decide whether to resume, suspend or roll back.

Freeze the loop and its impact

Build a per-instance timeline first. Automatic repair is a consequence; the first event is the health transition.

yaml vmss-repair-incident.yml
incident:
window_utc: 2026-09-12T08:00:00Z/2026-09-12T10:00:00Z
scale_set: vmss-workers-prod
expected_capacity: 20
useful_capacity: 15
health_source: application-health-extension
repair_action: Replace
grace_period: PT30M

capture_per_instance:
- instance_id
- model_version
- provisioning_time
- health_transitions
- repair_start_and_end
- application_release
- last_completed_work

hard_stops:
- useful-capacity-below-service-floor
- repair-removes-non-reconstructible-data
- repeated-replacement-of-the-same-capacity-slot

Set a capacity floor as well. Repairs are batched, but a pool that keeps recreating unhealthy instances can remain degraded. If headroom disappears, suspend automatic repairs as a tracked incident action before they consume the capacity needed for diagnosis.

Capture the policy and orchestration state

Read deployed configuration, not only the IaC template. Verify the action, grace period, health source and orchestration service state.

bash 01-capture-vmss-repair-state.sh
SUBSCRIPTION="<subscription-id>"
RG="rg-compute-prod"
VMSS="vmss-workers-prod"

az account set --subscription "$SUBSCRIPTION"

az vmss show \
--resource-group "$RG" \
--name "$VMSS" \
--query "{capacity:sku.capacity,upgradePolicy:upgradePolicy,automaticRepairsPolicy:automaticRepairsPolicy,orchestrationServices:orchestrationServices,extensions:virtualMachineProfile.extensionProfile.extensions}" \
--output json > vmss-repair-state.json

az vmss list-instances \
--resource-group "$RG" \
--name "$VMSS" \
--expand instanceView \
--query "[].{id:instanceId,latestModel:latestModelApplied,provisioning:provisioningState,health:instanceView.vmHealth.status.statuses[0].displayStatus}" \
--output table

A Suspended service state explains why an unhealthy instance is no longer repaired. Resuming it without understanding repeated failures only restarts the loop. An instance inside its grace period is not yet eligible for repair. That period begins after a state-changing operation completes and must cover real startup, warm-up and dependency registration.

Qualify the health signal

A VMSS uses either the Application Health extension or a Load Balancer health probe as the automatic repair health source. Prove which source is configured and replay its exact contract on one healthy and one affected instance.

text health-contract-review.txt
Application Health contract
Protocol, port and request path are explicit
Endpoint is reachable locally from the extension
Healthy response matches the selected binary or rich mode
Warm-up can report Initializing when rich states are enabled
Timeout is below the probe interval
Endpoint does not depend on a non-critical remote service

Negative controls
Stopped application becomes Unhealthy
Wrong path never reports Healthy
New instance remains non-eligible during initialization
Healthy status correlates with one useful workload transaction

With binary states, a response other than 200 or an unreachable endpoint becomes unhealthy. With rich health states, distinguish Initializing, Unknown and Unhealthy instead of treating every absence of Healthy as the same failure. A broken extension configuration and a genuinely failing application can emit the same operational symptom; local logs must separate them.

Prove what each repair action preserves

Replace, Reimage and Restart have different blast radii. Replace deletes the instance and creates one from the current VMSS model. Do not assume the instance ID, private IP or instance disks survive. Reimage preserves the instance and managed data disks but rebuilds the OS disk. Restart preserves disks and cannot repair a persistently broken model or OS.

yaml repair-action-contract.yml
Replace:
use_when: instance-is-disposable-and-current-model-is-known-good
verify: [externalized-state, ip-consumers, model-version, bootstrap-idempotence]

Reimage:
use_when: os-state-is-disposable-but-instance-identity-must-remain
verify: [data-disk-contract, bootstrap-after-reimage, extension-order]

Restart:
use_when: failure-is-transient-and-process-recovers-after-boot
verify: [restart-evidence, recurrence-window, service-autostart]

never_assume:
- local-state-survives-replace
- restart-fixes-model-drift
- reimage-preserves-os-disk-changes
- replacement-reuses-private-ip

If every replacement applies the same broken model, Replace is reproducing the incident. Compare model version, extensions, image, cloud-init and application settings between the last healthy instance and each replacement.

Separate grace period, bootstrap and model drift

Measure the real interval between completed provisioning and the first healthy signal. Do not extend the grace period from intuition alone.

text repair-loop-classification.txt
Healthy just after grace period expires
Grace period is shorter than measured startup and warm-up.

Never healthy, same error on every replacement
Current VMSS model, image, extension or bootstrap likely reproduces the fault.

Healthy then unhealthy under load
Diagnose application saturation, dependency failure or health endpoint design.

Healthy locally, Unknown in instance view
Diagnose Application Health extension configuration and reporting path.

One instance repeatedly unhealthy, model identical
Compare host, zone, attached resources and instance-specific initialization.

Keep the distribution, not only the average: P50, P95 and maximum readiness time across multiple creations. The grace period can then cover P95 with an explicit margin. Too long delays legitimate repair; too short turns normal startup into replacement.

Test one canary without restarting the fleet

Suspend repairs if required, correct one cause and create or reimage a single canary instance. Do not combine a new image, a new extension and a new repair action in the same test.

text vmss-repair-canary-gates.txt
Canary scope
One instance on the candidate model
Automatic repair remains suspended or isolated
Same health contract as production
No local state required for recovery

Promotion gates
Provisioning completes once
Health progresses through expected states
Readiness occurs inside the justified grace period
One useful workload transaction completes
Instance remains healthy across the recurrence window
Reapplying the model produces no drift

Refusal gates
Broken health path stays unhealthy
Failed bootstrap does not enter service
Capacity floor remains respected

The recurrence window must cover the observed trigger: first job, load peak, secret rotation or reboot. Two healthy minutes do not validate a failure that returned every hour.

Decide whether to resume, suspend or roll back

Change the repair action only when the failure mode and data contract justify it.

yaml vmss-repair-decision.yml
resume:
when:
- health-contract-is-proven
- current-model-creates-a-healthy-canary
- grace-period-covers-measured-readiness
- repair-action-matches-state-contract
action: resume-and-observe-one-repair-batch

suspend:
when:
- useful-capacity-is-at-risk
- service-state-was-suspended-after-repeated-failures
- replacement-reproduces-the-fault
action: preserve-instances-and-fix-model-first

rollback:
when:
- candidate-model-or-extension-breaks-health
- replacement-loses-required-state
- recurrence-window-fails
action:
- restore-last-known-good-model
- keep-automatic-repairs-suspended
- recreate-one-canary
- resume-only-after-positive-and-negative-tests

After resuming, observe a complete repair batch and verify that healthy instances return to expected capacity. Reconcile IaC afterward: a manual suspension or old image left outside code can recreate the loop during the next deployment.

Conclusion

A VMSS repair loop is not primarily a choice between Replace, Reimage and Restart. It is a disagreement between a health signal, startup time, machine model and persistence contract.

Resumption is safe when a canary built from the current model becomes healthy within a measured grace period, processes one useful transaction, fails the negative control correctly and survives the recurrence window. If replacement reproduces the fault or threatens capacity, suspend the policy, restore the last known good model and resume repairs only after the new canary passes validation.