Networking

Azure UDR: diagnose an NVA black hole before changing routes

A production runbook to separate effective routes, next hop, NVA health, IP forwarding, policy and return path before fixing a UDR or bypassing inspection.

13 Aug 2026 azureudrnvaroutingeffective-routesnetwork-watcherip-forwardinghub-spokeobservabilityautomationrunbookrollbackproduction

An application in an Azure spoke can no longer reach an internal API after a route-table change. DNS returns the expected address, no obvious NSG deny appears, and the 10.40.0.0/16 -> Virtual appliance route exists. The team is considering removing the UDR, opening the firewall or restarting the appliance. All three actions change the system before the loss point is known.

The running example is a TCP 443 flow from application VM 10.20.1.4 to service 10.40.2.10. The spoke sends traffic to an NVA in the hub at 10.10.0.4. This runbook drives an evidence-based decision: fix the UDR association or prefix, restore the appliance, repair the return path, roll back to the previous route or keep the change without bypassing inspection.

Freeze the exact flow and latest change

A “network outage” is too broad to test. Record the source, destination, protocol and port, then state the expected path in both directions. Preserve the first failure time and the removable change candidate without assuming that the latest change is already proven guilty.

yaml udr-incident-scope.yml
incident: inc-<id>
window_utc: <start>/<end>
source:
resource: vm-app-prod
nic: nic-vm-app-prod
ip: 10.20.1.4
subnet: snet-app-prod
destination:
ip: 10.40.2.10
port: 443
protocol: tcp
expected_path:
forward: spoke-app -> udr -> nva-hub -> service
return: service -> hub -> nva-hub -> spoke-app
change_candidate:
route_table: rt-spoke-app-prod
route: to-services-via-nva
deployment: <commit-or-change-id>
preserve:
- route table and subnet association before change
- effective routes on source and destination NICs
- Network Watcher next hop result
- NVA health, forwarding and session logs
- one failed probe with timestamp and correlation data

Freeze concurrent changes to route tables, peerings and the NVA during the initial capture. A diagnostic route can win through a more specific prefix and hide the symptom without explaining it.

Capture declared configuration before fixing it

The route table associated with a subnet is not yet proof of the route used by an interface. Capture the declared state, association and configured next hop before deleting anything.

bash 01-capture-udr-state.sh
set -euo pipefail

RG="<resource-group>"
VNET="<spoke-vnet>"
SUBNET="<source-subnet>"
ROUTE_TABLE="<route-table>"
NIC="<source-nic>"

az network route-table show --resource-group "$RG" --name "$ROUTE_TABLE" --output json > route-table-before.json

az network vnet subnet show --resource-group "$RG" --vnet-name "$VNET" --name "$SUBNET" --query '{id:id,addressPrefix:addressPrefix,addressPrefixes:addressPrefixes,routeTable:routeTable.id,networkSecurityGroup:networkSecurityGroup.id}' --output json > subnet-before.json

az network nic show --resource-group "$RG" --name "$NIC" --query '{id:id,ipConfigurations:ipConfigurations[].privateIPAddress,ipForwarding:enableIpForwarding}' --output json > source-nic-before.json

az network nic show-effective-route-table --resource-group "$RG" --name "$NIC" --output json > source-effective-routes-before.json

Review every prefix. A CIDR typo, a more specific route, a next hop still pointing to an old address or a table associated with the wrong subnet can each produce a black hole. Route names do not participate in route selection.

Prove the route Azure actually selects

Azure combines system routes, user-defined routes and propagated routes. Longest prefix wins first. For equal prefixes, a user-defined route takes precedence over a BGP route, which takes precedence over a system route. Read the Active or Invalid state, source, prefix and next hop in the source NIC effective routes.

Then run Network Watcher next hop for the exact destination. The expected result is VirtualAppliance, the intended IP and the responsible route table. A None result proves that no next hop exists. VirtualNetworkPeering, VirtualNetworkGateway or Internet when inspection is required proves that the flow is taking another path.

bash 02-prove-selected-next-hop.sh
set -euo pipefail

VM_RG="<source-vm-resource-group>"
VM="<source-vm>"
NIC="<source-nic>"
SOURCE_IP="10.20.1.4"
DEST_IP="10.40.2.10"

az network watcher show-next-hop --resource-group "$VM_RG" --vm "$VM" --nic "$NIC" --source-ip "$SOURCE_IP" --dest-ip "$DEST_IP" --output json > next-hop-source-to-destination.json

Run the test from the resource and NIC that actually originate the flow. An administration VM in another subnet can have a different effective route table and return a reassuring but irrelevant result.

Verify that the NVA can forward the packet

A correct next hop proves that the platform hands traffic to the appliance. It does not prove that the appliance forwards it. Check four layers independently: VM or scale-set state, enableIpForwarding on every transit NIC, guest forwarding, then the appliance policy and routing table.

text nva-forwarding-checks.txt
NVA control plan
Azure platform
  instance running with expected NIC attached
  next-hop private address still present
  enableIpForwarding=true on transit NICs
  healthy probe when a load balancer owns the next hop

Guest system
  IP forwarding enabled for the operating system and product
  expected interface, route and neighbor state available
  firewall or routing service running
  CPU, memory, sessions and queues not saturated

Policy
  rule allows source, destination, protocol and port
  NAT is applied only when the architecture requires it
  log shows receipt and forwarding of the same flow

Minimum evidence
  probe timestamp in UTC
  counters or logs on ingress and egress interfaces
  explicit reason for any drop
  identifier of the instance that handled the packet

When Azure NIC IP forwarding is disabled, also verify the equivalent guest or product setting before closing the diagnosis. When the next hop is an internal load balancer frontend, qualify its probes, HA Ports or specific rules and active backends. Do not replace that address with an individual appliance IP as an improvised fix.

Treat the return path as a second routing problem

A SYN on the NVA ingress without an application response does not prove a forward-path block. The destination may return through direct peering, a gateway or another appliance. A stateful firewall can then discard the asymmetric session, or the response can reach the source under a different identity.

Build a matrix with two independent route tests. For the forward direction, use the source NIC and destination IP. For the return direction, use a representative destination-side NIC and the source IP. Compare effective routes, route-table associations and appliance logs in the same time window.

text forward-return-evidence-matrix.txt
Forward 10.20.1.4 -> 10.40.2.10:443
selected effective route: <prefix/source/state>
observed next hop: <type/ip/route-table>
NVA receipt: <yes/no/timestamp>
NVA forwarding: <yes/no/reason>

Return 10.40.2.10 -> 10.20.1.4
selected effective route: <prefix/source/state>
observed next hop: <type/ip/route-table>
NVA receipt: <yes/no/timestamp>
NVA forwarding: <yes/no/reason>

Application
resolved target: <ip>
TCP connection: <success/timeout/reset>
TLS or HTTP response: <result>

Conclusion
<wrong-route|dead-next-hop|nva-drop|asymmetric-return|application>

This prevents the team from adding a return UDR by reflex. The right design depends on address domains, propagated routes and the inspection contract. A reciprocal route may be required, but it must be designed rather than improvised during the incident.

Separate routing, filtering and application failures

A timeout does not identify the layer that lost the flow. Classify the incident before changing it.

text udr-blackhole-classification.txt
Wrong route or association
Unexpected next hop, None, missing prefix or Invalid route
Action: fix the exact association, CIDR or next-hop address

Correct next hop, unavailable NVA
VirtualAppliance points to the right address but forwards no packets
Action: restore the NVA instance, backend or service

Incomplete forwarding
Packet arrives but is not forwarded; platform or guest forwarding is absent
Action: fix Azure and guest settings according to the NVA contract

Policy drop
NVA is healthy and explicitly denies the tested tuple
Action: make a targeted rule change with evidence and temporary expiry

Asymmetric return
Forward flow is visible; return is absent or uses another path
Action: restore symmetry or the intended NAT design

Application failure
TCP and TLS cross the NVA; application response fails
Action: remove route tables from the remediation scope

An NSG or firewall opening is justified only by filtering evidence. An incorrect effective route is not fixed by a security rule, and a healthy NVA is not repaired by removing inspection.

Choose one bounded correction

Prepare one reversible change. For a wrong route, correct the exact prefix or next hop and wait until the effective route shows the intended value. For a bad association, attach the correct table only to the affected subnet. For an unavailable NVA, restore its backend or fail over through the product’s designed HA mechanism.

Avoid three shortcuts: deleting a 0.0.0.0/0 route without inventorying affected flows, creating a temporary more-specific route without an expiry, or pointing directly to an instance behind a load balancer. Each can recover one probe while creating security or availability drift.

yaml bounded-routing-decision.yml
decision:
diagnosis: <wrong-route|dead-next-hop|forwarding|policy|asymmetric-return|application>
evidence_window_utc: <start>/<end>
owner: <operator>

change:
resource: <route-table|subnet|nva|policy>
exact_property: <prefix|next-hop|association|backend|rule>
previous_value: <captured-value>
proposed_value: <new-value>
blast_radius: <subnet-and-prefixes>
expires_at_utc: <timestamp-or-not-applicable>

validation:
- effective route is Active with expected source and prefix
- next hop matches the intended appliance path
- NVA sees forward and return packets
- TCP and application probe succeed from the real source subnet
- unrelated canary destinations keep their expected path

rollback:
trigger: <route-mismatch|probe-failure|asymmetry|unexpected-bypass>
action: restore captured property or previous IaC revision
post_check: replay route, next-hop and application probes

Validate and roll back without losing evidence

After the change, recapture the table, association, effective routes and next hop output. Replay the same probe from the same workload, to the same destination and port. Add a canary destination that should remain unaffected to detect an overly broad bypass.

Success requires five aligned proofs: expected active route, expected next hop, NVA transit in both directions, successful application probe and no canary regression. If one is missing, restore the captured value or previous IaC revision. Do not leave a diagnostic route outside code; it will become the next unexplained route.

Rollback should target the changed property. Restore one prefix, next hop or association instead of detaching the entire table to undo one route. After rollback, confirm that the faulty route disappeared from effective routes and that the previous path is visible again.

Conclusion

An NVA black hole behind an Azure UDR is a chain of evidence: associated table, selected effective route, reached next hop, forwarded packet, symmetric return and application response. A correct row in the declared route table validates only one link.

The production decision then becomes clear: fix the route when Azure selects the wrong path, restore the NVA when the packet stops there, repair symmetry when the return diverges, or remove networking from scope when the flow crosses inspection correctly. A sound rollback restores the exact property and replays the same evidence; it does not bypass inspection merely to remove the symptom.