Networking

Azure Network Watcher: diagnose an intermittent path before changing NSG or UDR

A production runbook for turning intermittent Azure connectivity into continuous evidence with Connection Monitor, then isolating DNS, routing, filtering and service health before changing network policy.

27 Sept 2026 azurenetwork-watcherconnection-monitorconnection-troubleshootnetworkingnsgudrdnsobservabilityincidentcanaryrunbookrollbackproduction

An application reaches a dependency most of the time, yet a few requests time out every hour. The NSG looks open, the route table has not changed and a manual test succeeds. Opening another rule or removing the UDR may hide the symptom, but it also destroys the evidence needed to distinguish a transient network fault from DNS rotation, backend saturation or an application timeout.

The running case is a VM-based order service in Azure calling orders-db.internal.example over TCP 5432 through a hub firewall. Failures affect only some instances and last less than two minutes. This runbook uses Azure Network Watcher Connection Monitor to preserve the time series, then Connection Troubleshoot and data-plane evidence to isolate the failing layer. The outcome is an explicit decision to correct one bounded component, hold the change or restore the previous path.

Freeze the flow contract before testing

Do not begin with a generic ping. Record the application flow exactly as production uses it: source cohort, destination FQDN, resolved addresses, protocol, port, expected path, timeout budget and the last known good window. Include every source instance that can serve traffic; one healthy VM does not clear a scaled service.

yaml 01-intermittent-flow-contract.yml
incident: INC-NET-427
window_utc: 2026-09-27T05:40:00Z/2026-09-27T06:20:00Z
sources:
resource_group: rg-orders-prod
instances: [vm-orders-01, vm-orders-02, vm-orders-03]
destination:
fqdn: orders-db.internal.example
protocol: TCP
port: 5432
expected_path:
next_hop: azure-firewall
egress_identity: firewall-policy-prod
service_budget:
connection_timeout_ms: 2000
failed_checks_percent: less-than-1
changes_to_freeze:
- DNS records and TTL
- subnet NSG and security admin rules
- UDR and route propagation
- firewall policy
- backend listener and deployment

Capture the Connection Monitor resource and the affected network configuration before editing them. Keep the monitor definition, endpoint membership, test groups, protocol, frequency and success thresholds. A monitor that tests the wrong source subset or only an IP while the application uses an FQDN can be healthy and still fail to represent the incident.

bash 02-freeze-network-evidence.sh
SUB="<subscription-id>"
RG="NetworkWatcherRG"
LOCATION="westeurope"
WATCHER="NetworkWatcher_westeurope"
MONITOR="orders-to-db-prod"
CM_ID="/subscriptions/$SUB/resourceGroups/$RG/providers/Microsoft.Network/networkWatchers/$WATCHER/connectionMonitors/$MONITOR"

az resource show --ids "$CM_ID" --output json > connection-monitor-before.json

az network nic show-effective-route-table   --resource-group rg-orders-prod   --name nic-vm-orders-02   --output json > effective-routes-before.json

az network nic list-effective-nsg   --resource-group rg-orders-prod   --name nic-vm-orders-02   --output json > effective-nsg-before.json

The files are rollback inputs and incident evidence. They do not prove the path is healthy.

Make the continuous test represent production

Connection Monitor organizes endpoints, test configurations and test groups. At the individual test level, it records failed-check percentage and round-trip time. Design the test around an operational question, not around resource inventory.

For this case, use every serving VM as a source, the production FQDN as the destination and TCP 5432 as the test. Keep DNS resolution in the path by monitoring the FQDN rather than pinning the current address. If the application also depends on TLS or an HTTP readiness contract, add a separate test configuration; a TCP handshake cannot prove an HTTP response, authentication or useful database query.

Confirm that the sources are actually reporting. Missing test data is not successful connectivity. Separate these states:

  • checks succeed within the latency budget;
  • checks fail or latency rises;
  • one source stops reporting;
  • the test resolves or targets a different endpoint than the application.

Set alerts on failed-check percentage and round-trip time with a window long enough to avoid a single probe becoming an incident, but short enough to capture the reported burst. Alert separately on absence of monitoring data when your monitoring design supports it. A single aggregate across all sources can dilute a failure limited to one VM, so inspect the source-destination-test combination.

Correlate loss with the affected cohort

Start with the Connection Monitor timeline. Compare failure percentage and RTT per source over the frozen UTC window. The shape narrows the hypothesis:

  • one source fails while peers stay healthy: inspect its NIC, effective rules, routes, guest firewall and host pressure;
  • every source fails toward one resolved address: inspect DNS rotation, the destination instance and its return path;
  • all destinations through the same next hop degrade: inspect the hub firewall, NVA, gateway or route propagation;
  • RTT rises before loss: inspect saturation and queueing before opening security rules;
  • the monitor stays healthy while application calls fail: move up to TLS, authentication, connection pooling or the application timeout.

Do not average away the incident. Preserve the smallest failing source, destination address and interval. That tuple becomes the input to the on-demand diagnostic and log queries.

Run Connection Troubleshoot during the failure

Connection Monitor proves when and where the symptom repeats. Connection Troubleshoot is the bounded diagnostic for the exact tuple. Run it from the affected source to the FQDN or address and port while the problem is active. Retain connectivity status, probes sent and failed, latency, hops, next-hop analysis and reported issues.

The result can expose DNS resolution, an NSG decision, a user-defined route, a guest firewall, a port with no listener or resource pressure. Treat it as a snapshot, not as a replacement for the continuous series. A successful run after the burst only proves the path at that later timestamp.

Cross-check each conclusion with a focused tool:

  • use Next Hop to prove the effective route for the resolved destination;
  • use IP Flow Verify or NSG diagnostics to identify the deciding security rule;
  • inspect effective routes and rules on the affected NIC;
  • read VNet flow logs or firewall logs for the same five-tuple and UTC window;
  • validate the destination listener and application logs.

An Allowed NSG result does not prove the firewall allowed the session, and a correct next hop does not prove the return path. Likewise, an absent flow-log record may mean collection did not cover the flow; it is not automatically proof of a drop.

Separate DNS rotation from path instability

An intermittent FQDN often resolves to several addresses. Record answers from each affected source repeatedly during the window and compare them with the monitor’s destination evidence. Check CNAME chain, TTL, resolver path and whether every returned address belongs to the approved backend set.

bash 03-sample-dns-and-tcp.sh
TARGET="orders-db.internal.example"
PORT="5432"

for attempt in 1 2 3 4 5; do
date -u +"%Y-%m-%dT%H:%M:%SZ"
getent ahostsv4 "$TARGET" | awk '{print $1}' | sort -u
timeout 3 bash -c "true >/dev/tcp/$TARGET/$PORT"     && echo "tcp=ok"     || echo "tcp=failed"
sleep 10
done

Run this from a representative production source or an approved diagnostic host on the same path. If one address fails consistently, keep the FQDN contract and repair or remove the unhealthy backend through its owning service. Do not pin a hosts-file entry as the production fix. If answers differ by source, investigate resolver forwarding, cache and split-horizon policy before changing the NSG or UDR.

Prove forward and return paths

For routed hub-and-spoke or hybrid traffic, validate both directions. The source route can correctly select Azure Firewall while the destination returns through another hub, VPN, ExpressRoute path or local default. Stateful devices then see only half the conversation.

Correlate a failed check with firewall allow or deny records and, where available, session creation and teardown. Verify SNAT behavior and the source identity observed by the destination. Compare effective routes on both endpoint subnets, including propagated routes and more-specific prefixes.

If the evidence points to an NSG or UDR, name the exact rule or prefix. “Networking issue” is not an actionable diagnosis. The correction should be as narrow as the evidence: one priority, one prefix, one next hop, one backend route or one return-path advertisement.

Canary one correction without erasing the baseline

Keep the original monitor active. Apply the candidate to one source subnet, one route prefix, one firewall rule collection or one backend, depending on the proven fault. Record the change identifier and start time so the metric series can separate before and after.

Use three controls:

  1. The positive flow must recover from the affected source and stay within its loss and latency budget.
  2. A peer source on the unchanged path must remain stable.
  3. A destination that should be denied must stay denied.

Observe several full test intervals and an application transaction, not only one successful probe. Watch the same source-destination series, firewall evidence and backend logs. If the candidate only moves the failure to another source or causes the negative control to pass, stop and roll it back.

Decide, validate or roll back

Promote the correction when Connection Monitor stays within budget across every required source, Connection Troubleshoot no longer reports the diagnosed issue, the application transaction succeeds and the refusal control remains enforced. Capture the final effective routes, rules and monitor definition in the change record.

Hold when the monitor detects the symptom but the failing layer remains ambiguous. Continue evidence collection rather than widening network access. A short intermittent failure with no correlated data is a reason to improve observability, not to remove controls.

Roll back when loss or latency worsens, the next hop deviates from the contract, an unintended flow becomes reachable, or application health does not recover. Restore the saved IaC revision or the exact rule, route or monitor configuration changed by the canary. Then prove that the prior path and monitoring coverage are back; a configuration deployment alone is not a completed rollback.

Conclusion

Intermittent connectivity is difficult because manual success is easy to obtain after the evidence has disappeared. Connection Monitor turns the path into a source-specific time series, while Connection Troubleshoot, effective configuration and traffic logs explain one failing interval.

The production decision follows the proof: correct the DNS member, route, security rule, firewall path or backend that actually failed; hold when the layer is still uncertain; or restore the previous configuration when the canary breaks the contract. The objective is not to make the next probe green. It is to make the path continuously explainable and reversible.