Networking

Azure Virtual WAN: validate Routing Intent before enforcing production inspection

A production runbook for qualifying Azure Virtual WAN Routing Intent across private flows, internet egress, effective routes, Azure Firewall, canaries, validation and rollback.

20 Sept 2026 azurevirtual-wanrouting-intentazure-firewallnetworkingroutingexpressroutevpnobservabilitykqlcanaryrunbookrollbackproduction

A network team wants to enforce east-west and internet traffic inspection in a secured Azure Virtual WAN hub. Enabling Routing Intent looks simpler than maintaining route-table associations, propagations and static routes for every connection. Minutes after the change, an application in one spoke can no longer reach an on-premises API, another spoke still has internet access, and Azure Firewall logs do not contain every expected flow.

The running case is a Standard Virtual WAN with several hubs, VNet connections, site-to-site VPN, ExpressRoute and Azure Firewall in each secured hub. The runbook must decide whether private and internet traffic policies can be enabled, whether the rollout should remain on one hub, or whether the previous routing state must be restored. A green provisioning state is not enough: every traffic family needs a proven forward and return path.

Freeze the traffic contract before the policies

Routing Intent controls two traffic families. The Private Traffic policy steers traffic between branches, VNets and hubs through the selected security solution. The Internet Traffic policy makes eligible connections learn a default route and sends internet egress through that next hop. A mistake in the first policy breaks internal reachability. A mistake in the second can move every outbound dependency of a workload.

Start with the matrix that must remain true after cutover.

yaml routing-intent-contract.yml
change: vwan-routing-intent-weu-prod
hub: vhub-weu-prod
security_next_hop: azfw-vhub-weu-prod

flows:
- name: spoke-to-onprem-api
  source: 10.42.8.0/24
  destination: 172.20.40.15:443
  policy: private
  expected: allow-and-log
- name: onprem-to-spoke-sql
  source: 172.20.12.0/24
  destination: 10.42.20.4:1433
  policy: private
  expected: allow-and-log
- name: spoke-to-spoke-health
  source: 10.42.8.0/24
  destination: 10.43.6.10:443
  policy: private
  expected: allow-and-log
- name: spoke-to-internet
  source: 10.42.8.0/24
  destination: approved-egress-target:443
  policy: internet
  expected: allow-log-and-stable-egress-ip

stop_conditions:
- missing_reverse_path
- firewall_does_not_observe_expected_flow
- default_route_learned_by_unapproved_connection
- private_prefix_missing_from_effective_routes

rollback: redeploy-snapshot-vwan-routing-v17

Include flows that must remain denied. A positive probe without a negative probe only proves that a path is open, not that inspection enforces the intended policy.

Inventory what Routing Intent will replace

Routing Intent manages the hub connections’ route-table associations and propagations. It cannot be introduced into a hub that still depends on custom route tables or static routes whose next hop is a VNet connection. Before enabling it, export the complete state: hubs, connections, tables, labels, static routes, propagations, gateways, peerings, BGP prefixes and the default-route propagation setting.

Do not reduce this step to screenshots. The rollback artifact must be deployable and versioned. Removing Routing Intent later does not automatically reconstruct the former defaultRouteTable configuration.

Use Azure Resource Graph to confirm the exact target hub and preserve resource identities.

kusto 01-inventory-vwan-routing-intent.arg.kql
resources
| where type =~ "microsoft.network/virtualhubs/routingintent"
 or type =~ "microsoft.network/virtualhubs/hubvirtualnetworkconnections"
 or type =~ "microsoft.network/virtualhubs/hubroutetables"
| extend hubId = tostring(split(id, "/routingIntent/")[0])
| project subscriptionId, resourceGroup, name, type, location,
        hubId, provisioningState=tostring(properties.provisioningState), properties
| order by resourceGroup asc, type asc, name asc

The snapshot must also record the current internet access model. Direct egress through Azure Firewall has a different contract from forced tunneling through an on-premises site or NVA. Preserve observed public egress IPs and the dependencies that allowlist them.

Prove prerequisites without confusing control and data planes

The hub must be eligible, and the security solution must be ready. For Azure Firewall, verify at least its provisioning state, attached Firewall Policy, capacity, diagnostics, rules required by the canaries and intended SNAT behavior. For an integrated NVA or SaaS solution, add instance health, internal and external interfaces, appliance-local routes and vendor limits.

Keep three separate proofs:

  • the Azure control plane accepts the configuration;
  • effective routes identify the expected next hop;
  • packets actually cross the security solution and return on a coherent path.

A Succeeded deployment proves only the first. A visible route does not prove that the Firewall Policy allows the flow, SNAT is correct, or the destination can reply.

Read effective routes from both ends

After applying the candidate in preproduction or on the first bounded hub, inspect the effective routes on the private-policy next hop. When only an internet policy is enabled, inspect the defaultRouteTable as well. For every prefix in the contract, verify its origin, next hop and state.

From a canary VM in each traffic family, also inspect NIC effective routes and use Network Watcher Next Hop or Connection Troubleshoot against the exact destination. Repeat from the opposite direction. For VNet-to-on-premises traffic, a correct spoke route does not prove that the ExpressRoute or VPN return path crosses the same hub and firewall.

In a multi-hub topology, enable and validate the private policy consistently on every hub that must inspect inter-hub traffic. A partial rollout can create an inspected forward path and a return path that prefers another hub. Freeze the hub routing preference and competing VPN, ExpressRoute or NVA advertisements before blaming the firewall.

Handle non-RFC 1918 private prefixes explicitly

Routing Intent naturally recognizes the expected private address spaces, but some enterprises carry internal non-RFC 1918 prefixes. Declare those ranges as additional private prefixes on every relevant hub. A declaration on one hub does not automatically propagate to the others.

For each additional range, verify two effects: the private policy steers it correctly, and the security solution does not apply unwanted SNAT. A connection may work while becoming operationally useless because the destination no longer sees the expected source address.

Keep a compact register with prefix, owner, origin hub, hubs that must learn it, expected policy, SNAT behavior and validation probe. Any range without an owner or canary blocks promotion.

Separate direct access from forced tunneling

Choose the internet model before enabling the policy.

With direct access, the internet policy steers flows to the hub security component, which inspects them and sends them to the internet. Verify that every connection expected to learn 0.0.0.0/0 has default-route propagation enabled, required destinations are allowed, and observed egress IPs match the contract.

With forced tunneling, the private policy also carries 0.0.0.0/0 as an additional prefix, and the security next hop forwards traffic through a locally learned default route from ExpressRoute, VPN, an NVA or a supported static route. The default route does not propagate across hubs, so every hub needs a valid local source. Without one, internet traffic is blocked after inspection.

Do not mix both models in one cutover. Explicitly test DNS, secret retrieval, container registries, license activation, monitoring endpoints and package repositories. These quiet dependencies often expose a poorly qualified default route or SNAT policy first.

Correlate probes, routes and Firewall logs

Run canaries with an identifier, a short UTC window and a stable source. Evidence must connect the test, effective route, firewall rule and application result.

kusto 02-correlate-routing-intent-canaries.kql
let StartUtc = datetime(<start-utc>);
let EndUtc = datetime(<end-utc>);
let CanarySources = dynamic(["10.42.8.4", "10.43.6.4"]);
union isfuzzy=true AZFWNetworkRule, AZFWApplicationRule
| where TimeGenerated between (StartUtc .. EndUtc)
| where SourceIp in (CanarySources)
| project TimeGenerated, SourceIp, DestinationIp, DestinationPort,
        Fqdn, Protocol, Action, RuleCollectionGroup, RuleCollection, Rule
| order by TimeGenerated asc

Adapt table and column names to the workspace diagnostic mode. A missing log is not automatically a deny: the flow might not reach the firewall, collection might be late, or the expected category might be disabled. Compare with Firewall metrics, Network Watcher and the target log.

Require four results for every row in the matrix: expected next hop, expected inspection log, expected application response and proven return path. Measure latency before and after as well. Inspection that works but breaches the latency budget is still an invalid cutover.

Roll out hub by hub inside a bounded window

Routing Intent applies to the hub, not to one canary spoke. Risk reduction therefore starts with a representative preproduction hub and continues with the production hub that has the smallest blast radius. Prepare probes before the change window and run them as soon as routes converge.

Freeze concurrent changes to BGP, hub connections, route tables and Firewall Policy during cutover. Otherwise a broken flow can be attributed to the wrong change.

The observation window must cover both persistent and new connections. An established socket can survive while a new DNS resolution or route fails. Force fresh sessions and include private traffic, internet traffic, an expected deny and an inter-hub return path.

Decide promotion, hold or rollback

Promote when every contract prefix has the correct next hop, forward and return canaries pass, denies remain effective, Firewall logs explain each decision, egress retains its intended identity, and latency stays within budget. Move to the next hub with the same evidence set.

Hold on the canary hub when routes have converged but the return path, SNAT behavior, a non-RFC 1918 prefix or observability remains ambiguous. Do not widen the rollout to collect more signal. Narrow the experiment to one destination and one advertisement.

Roll back when a private prefix disappears, an unintended connection learns the default route, the path becomes asymmetric, the firewall does not observe a flow it must inspect, or a critical dependency loses egress. Rollback means redeploying the versioned snapshot of associations, propagations and routes, then replaying the canaries. Deleting Routing Intent without restoring that state is not a return plan.

Conclusion

Routing Intent simplifies operations only when the declarative intent remains verifiable in the data plane. The useful validation unit is not a policy or hub in isolation. It is a named flow with its forward route, inspection, return path and application evidence.

The production decision is then explicit: promote hub by hub with effective routes and complete canaries, hold while any prefix or SNAT behavior remains ambiguous, and restore the versioned state at the first unexplained path. Centralized routing gains the property it needs most in production: reversibility.