Networking
Azure Traffic Manager: diagnose a degraded profile before forcing failover
A production runbook for separating application health, probe contract, endpoint state, DNS answers, TTL and secondary capacity before forcing Traffic Manager failover.
An Azure Traffic Manager profile changes to Degraded while the primary endpoint still answers the team’s checks. Some users reach the secondary region, others keep the old destination, and pressure builds to disable the primary endpoint manually. That move can amplify the incident when the probe targets the wrong path, DNS caches have not expired, or the secondary region cannot absorb production load.
This runbook supports an evidence-based decision: repair the probe or backend, let automatic failover proceed, force a bounded cutover, or restore the previous configuration. Traffic Manager distributes DNS answers. It does not move established connections, and it cannot make every downstream resolver adopt a new answer immediately.
Freeze the failover contract
Write down the intended behavior first. A degraded status alone does not tell you which endpoint receives traffic, why a probe fails, or how many clients still hold a cached answer.
profile: tm-orders-prod
resource_group: rg-global-routing-prod
routing_method: Priority
public_name: api.example.com
traffic_manager_name: orders-prod.trafficmanager.net
incident_window: 2026-09-08T15:10:00Z/2026-09-08T15:40:00Z
endpoints:
primary: app-orders-weu.example.net
secondary: app-orders-neu.example.net
probe_contract:
protocol: HTTPS
port: 443
path: /health/ready
expected_status: 200-299
host_header: api.example.com
decision:
failover_only_if: primary-unhealthy-and-secondary-capacity-proven
rollback_on: secondary-errors-or-dns-answer-mismatch Retain the IaC commit, last probe change, priorities, TTL and exact incident window. Name the team that validates secondary application capacity as well. An Online endpoint proves a health response, not a load budget.
Read the deployed profile and endpoint state
Separate administrative state from monitor state. An endpoint can be Enabled and still be Degraded. AlwaysServe can also keep it in the routing method without health evaluation. Read Azure before changing the portal.
RG="rg-global-routing-prod"
PROFILE="tm-orders-prod"
az network traffic-manager profile show \
--resource-group "$RG" \
--name "$PROFILE" \
--query '{status:profileStatus,monitor:monitorConfig,dns:dnsConfig,method:trafficRoutingMethod}' \
--output jsonc
az network traffic-manager endpoint list \
--resource-group "$RG" \
--profile-name "$PROFILE" \
--query '[].{name:name,type:type,status:endpointStatus,monitor:endpointMonitorStatus,target:target,priority:priority,weight:weight,alwaysServe:alwaysServe}' \
--output table Check endpoint type, target, priority or weight, and any nested profiles. A forced change on the parent profile will not repair an inconsistent minChildEndpoints threshold in a child profile.
Replay the exact probe contract
A health probe is not a complete user request. It uses the configured protocol, port, path, status code ranges and headers. A redirect, certificate mismatch, missing host header, authentication challenge or slow dependency can fail the probe while the home page still works from an administrator’s machine.
PRIMARY="app-orders-weu.example.net"
SECONDARY="app-orders-neu.example.net"
HOST="api.example.com"
PATH_READY="/health/ready"
for target in "$PRIMARY" "$SECONDARY"; do
curl --silent --show-error --output /dev/null \
--write-out "$target status=%{http_code} connect=%{time_connect}s total=%{time_total}s\n" \
--header "Host: $HOST" \
--max-time 10 \
"https://$target$PATH_READY"
done Replay from several authorized external networks. Do not use -k as health evidence: bypassing TLS validation hides one of the likely causes. When the endpoint relies on SNI and the public hostname, use a controlled curl --resolve test instead of calling the IP address directly.
The probe should test useful but bounded readiness. A check that traverses every dependency can withdraw a region for a noncritical failure. A probe that always returns 200 can keep routing to an endpoint that cannot serve real requests.
Read probe events, not only current status
The current status hides the timeline. The ProbeHealthStatusEvents resource logs identify the endpoint, transitions and result description. Correlate them with backend logs over the same UTC window.
let Start = datetime(2026-09-08T15:00:00Z);
let End = datetime(2026-09-08T15:50:00Z);
AzureDiagnostics
| where TimeGenerated between (Start .. End)
| where ResourceType == "TRAFFICMANAGERPROFILES"
| where Category == "ProbeHealthStatusEvents"
| project TimeGenerated,
EndpointName = EndpointName_s,
ProbeStatus = Status_s,
ResultDescription,
ResourceId = _ResourceId
| order by TimeGenerated asc One transition followed by a quick recovery can be transient. Persistent failures aligned with backend 5xx, TLS timeouts or saturation justify action on the region. Missing logs mean diagnostic settings must be checked first; they do not prove that probes succeeded.
Prove the DNS answer clients receive
Traffic Manager works at the DNS layer. Query several resolvers and retain the CNAME chain, final answer and TTL. Repeating a local lookup can merely read the corporate cache again.
NAME="api.example.com"
for resolver in 1.1.1.1 8.8.8.8; do
echo "resolver=$resolver"
dig +noall +answer @$resolver "$NAME" CNAME
dig +noall +answer @$resolver "$NAME" A
done
# Query the corporate resolver from a real user network as well.
dig +noall +answer "$NAME" CNAME
dig +noall +answer "$NAME" A An endpoint state change does not terminate existing connections. Resolvers and clients can retain the previous answer until TTL expiry, and sometimes beyond it depending on their behavior. Measure convergence from multiple viewpoints instead of declaring failover complete as soon as the portal changes.
Prove the secondary region can carry production
Before disabling the primary endpoint, validate the secondary as a production target, not merely as a healthy URL.
Secondary capacity
Sufficient instances, quotas and autoscaling
Regional dependencies available
Connection pools and downstream limits understood
Application contract
Same or compatible release
Expected configuration and feature flags
Replication lag measured against the RPO
Writes and asynchronous work can be reconciled
User path
Certificate and host header valid
WAF, rate limits and allowlists aligned
Synthetic read then bounded write validated
Telemetry and alerts active in the secondary region The Endpoint Status by Endpoint metric confirms the state observed by Traffic Manager. Queries by Endpoint Returned, split by endpoint, shows the shift in DNS answers. It measures neither active connections nor HTTP requests actually served, so pair it with application and dependency metrics.
Exercise failover without creating a lasting outage
For a planned exercise, lower TTL early enough for the previous value to expire before the test. Lowering it during an incident does not erase answers already cached. Prefer a dedicated canary hostname or validation endpoint when the routing design supports one.
preconditions:
- secondary-readiness-gate-passed
- probe-events-correlated-with-backend
- current-dns-answer-and-ttl-captured
- rollback-owner-present
canary:
scope: controlled-client-group
verify:
- dns-answer-points-to-secondary
- tls-and-host-header-valid
- synthetic-read-and-bounded-write-succeed
- secondary-error-rate-and-latency-within-budget
stop_conditions:
- secondary-capacity-saturation
- replication-lag-outside-rpo
- rising-5xx-or-dependency-errors
- unexpected-dns-answer Do not enable AlwaysServe to clear a red status. It bypasses the health decision and can return an unhealthy endpoint. Avoid changing probe, priority, TTL and backend configuration together; doing so destroys attribution.
Decide, validate or roll back
fix_probe_or_backend:
when:
- deployed-probe-does-not-match-readiness-contract
- logs-prove-tls-host-header-status-or-timeout-failure
validate:
- endpoint-returns-online
- user-path-and-probe-both-pass
allow_automatic_failover:
when:
- primary-failure-is-proven
- secondary-capacity-and-data-readiness-are-proven
- dns-ttl-convergence-is-observed
force_bounded_failover:
when:
- primary-still-receives-harmful-traffic
- automatic-state-does-not-match-proven-failure
- change-owner-stop-conditions-and-rollback-are-ready
rollback:
action:
- restore-endpoint-status-priority-and-probe-from-iac
- reenable-primary-only-after-readiness-passes
- observe-dns-answers-for-at-least-one-ttl-window
- reconcile-writes-and-async-work-in-both-regions After cutover, validate three planes separately: DNS answers return the intended target, that target serves real requests, and business effects remain consistent. If the secondary region degrades, return to the last known IaC state and restore the primary only after its readiness is proven. Oscillation between regions is harder to recover from than a delayed failover.
Conclusion
A Degraded Traffic Manager profile is a diagnostic signal, not an instruction to force a manual cutover. The decision must connect deployed configuration, probe contract, health events, DNS answers, TTL, secondary capacity and data state.
Repair the probe when it measures the wrong contract, repair the backend when failure is real, let automatic failover work when its preconditions are proven, and force a cutover only with stop conditions and rollback. Final validation belongs to DNS, the application and business effects together.