Infrastructure
Azure Backup: validate a VM restore before declaring the recovery path ready
A production runbook for testing an Azure VM recovery point in isolation, validating the system and application, measuring recovery, then deciding readiness, remediation or rollback.
An Azure Backup job marked Completed proves that a recovery point was produced. It does not prove that the team can select the right point, restore the disks, boot the VM without reconnecting a duplicate to production, and validate the application within the expected time.
The running example is an internal API hosted on an Azure VM with an OS disk and a data disk. The team is preparing a recovery drill after an image, network or backup policy change. The runbook must end with an explicit decision: recovery path ready, conditionally ready, or not ready with the previous configuration restored.
Write the restore contract before selecting a point
A drill does not start in the portal. It starts with the workload to recover, acceptable data loss, business checks and effects the restored VM must never produce.
workload: api-orders-prod
source_vm: vm-orders-01
vault: rsv-platform-prod
protected_item: vm-orders-01
restore_mode: alternate_location
target_resource_group: rg-restore-drill-202609
target_network: vnet-restore-isolated
recovery_point_rule: newest_point_before_change_window
required_consistency: application_or_documented_recovery_procedure
business_checks:
- database-opens-read-only
- last-approved-order-present
- api-health-contract-valid
forbidden_effects:
- production-dns-registration
- outbound-email-or-webhook
- queue-consumption
- write-to-production-dependency
cleanup_owner: platform-operations
rollback: delete-drill-resources-and-keep-evidence RPO and RTO must be decision criteria, not numbers copied from a policy. Verify RPO in the data that was actually restored. Measure RTO from the start of the exercise to business validation, including point selection, restore, boot and checks.
Inventory the vault, protected item and available points
Freeze identifiers so the drill cannot silently test a namesake VM, another vault or a point created after the simulated incident. These commands create a reproducible inventory; use native names when multiple protected items share the same friendly name.
VAULT_RG="rg-backup-prod"
VAULT="rsv-platform-prod"
VM="vm-orders-01"
az backup item list --resource-group "$VAULT_RG" --vault-name "$VAULT" --backup-management-type AzureIaasVM --query "[?properties.friendlyName=='$VM'].{name:name,health:properties.healthStatus,last:properties.lastBackupTime,policy:properties.policyId}" --output json
az backup recoverypoint list --resource-group "$VAULT_RG" --vault-name "$VAULT" --backup-management-type AzureIaasVM --container-name "$VM" --item-name "$VM" --query "[].{name:name,time:properties.recoveryPointTime,type:properties.recoveryPointType}" --output table Retain the simulated incident time and selected recovery point time in UTC. “Latest” is not always correct: a point created after logical corruption may already contain the faulty state.
Qualify consistency instead of assuming it
Azure VM Backup can produce recovery points with different consistency characteristics depending on the OS, configuration and application processing result. Point availability therefore does not replace reading its type or testing the application’s recovery procedure.
For every candidate point, record:
- its UTC time and observed consistency type;
- included disks and their purpose;
- the associated backup job result;
- application recovery steps required after a crash-consistent start;
- why it fits the change and incident window.
If the application requires transactional consistency, a successful boot is not validation. The database must complete recovery, open the expected structures and return a coherent business sample. Without that check, the path is conditionally ready at best.
Isolate the target before any restore
Alternate-location restore protects the source from replacement, but it may create a second system with the same hostnames, schedules, identities and dependency configuration. Isolation must be effective before the first boot.
Surface Control before boot
Network Dedicated VNet and subnet, no peering or route to production
Egress Deny by default; allow only explicit test targets
DNS Explicit test resolution, no automatic production registration
Identity No implicit write access to production resources
Messaging Consumers, webhooks, SMTP and notifications disabled or redirected
Scheduling Cron, timers and deployment agents neutralized before validation
Addressing No production name or IP reused
Observability Drill logs separated and correlated with an exercise ID A Private Endpoint may participate in a restore path for some services, but it is not the main diagnostic lens. The broader question is whether the isolated target can retrieve what it needs without gaining a production write path.
Restore disks to an alternate location
Restore disks into a dedicated resource group first. This leaves a control point before a VM is created and started. Exact arguments depend on the vault and protected item; use identifiers captured during inventory.
VAULT_RG="rg-backup-prod"
VAULT="rsv-platform-prod"
CONTAINER="vm-orders-01"
ITEM="vm-orders-01"
RECOVERY_POINT="<recovery-point-name>"
STAGING_ACCOUNT="strestore202609"
TARGET_RG="rg-restore-drill-202609"
az backup restore restore-disks --resource-group "$VAULT_RG" --vault-name "$VAULT" --container-name "$CONTAINER" --item-name "$ITEM" --rp-name "$RECOVERY_POINT" --storage-account "$STAGING_ACCOUNT" --target-resource-group "$TARGET_RG" --restore-mode AlternateLocation Capture the returned job name and follow it with az backup job show or az backup job wait. Do not move to VM creation until the job is complete and restored artifacts match the expected disk set.
Staging storage, permissions, encryption, zones and special VM configurations can change the procedure. The local runbook must document those prerequisites instead of discovering them during an incident.
Build the VM without reimporting production effects
Inspect the restored template and disks before deployment. Create the VM in the drill resource group and subnet, without a public IP, after egress controls are in place. Replace or disable anything that may cause effects: deployment extensions, startup scripts, monitoring agents using production identifiers and scheduled tasks.
The expected result is more than PowerState/running. Prove three layers:
- the operating system boots and volumes mount in the expected locations;
- the application reaches a stable state without production dependencies;
- restored data satisfies the business contract and chosen point window.
Run technical and business validation
Use the same evidence manifest for every exercise. It prevents one operator from closing on a ping while another expects a complete business transaction.
drill_id: restore-20260911-orders
recovery_point_utc: "<timestamp>"
restore_job_status: completed
checks:
os_boot: pass
expected_disks_mounted: pass
filesystem_check: pass
application_started: pass
health_endpoint_contract: pass
database_recovery: pass
business_sample: pass
production_write_refusal: pass
outbound_side_effect_refusal: pass
timing:
exercise_started_utc: "<timestamp>"
restore_completed_utc: "<timestamp>"
business_validation_completed_utc: "<timestamp>"
decision: pending The refusal test matters. A VM that reads the right data but can also consume the production queue is not a safe drill target. Check routes, network rules, effective permissions and denial logs as carefully as the positive test.
Measure the complete path and process gaps
Separate Azure service time, operator wait time and application validation time. This breakdown makes improvement actionable: automating protected-item selection has little value when manual VM reconstruction consumes most of the recovery window.
Record every deviation from the source as well: unavailable VM size, missing extension, expired secret, unresolved DNS dependency, non-automated recovery procedure or incomplete business check. A restore that succeeded after three improvised fixes does not prove a ready runbook; it provides the remediation backlog.
Decide readiness, remediation or rollback
READY
Recovery point selected against the simulated incident window
Restore completed from documented identities and commands
Isolated VM booted with every expected disk
Technical and business checks passed
Forbidden production effects were refused
Measured recovery fits the approved objective
Cleanup and next-test owner are recorded
READY WITH CONDITIONS
Data is valid but one manual step remains documented
Recovery objective is missed with an accepted remediation plan
One non-critical dependency requires a controlled workaround
NOT READY
No suitable recovery point exists
Point consistency does not satisfy the application contract
Restored target can write to production
Data or business validation fails
Restore cannot be reproduced from the runbook
ROLLBACK THE RECENT CHANGE
Readiness regressed after a policy, vault, identity or network change
Restore the previous known configuration
Repeat the drill before closing the change Drill rollback means deleting temporary resources after evidence is retained. It does not mean changing the source VM or disabling protection. If the test exposes a regression caused by a recent change, restore the known configuration and repeat the entire path. A reverted configuration without a new drill remains an assumption.
Conclusion
An operable backup is judged when a team restores, isolates, boots and validates the workload, not by the number of green jobs in the vault. The useful runbook connects the selected point to a simulated incident, qualifies consistency, prevents side effects and measures time until business proof.
The decision is then clear: declare the path ready when restore, refusal and validation are reproducible; accept a documented condition temporarily; or roll back the recent configuration and retest. Azure Backup supplies the recovery point. The drill turns it into demonstrated recovery capability.