Networking
Azure Front Door: diagnose an unhealthy origin before changing routing
A production runbook to separate health probes, DNS/TLS, Private Link, backend capacity and routing before draining, failing over or rolling back.
An API delivered through Azure Front Door starts returning 503 responses in one region. One origin is unhealthy and traffic shifts to the second backend. The immediate response is often to set its weight to zero, increase probe tolerance or disable probing. Those changes might restore service, but they can also hide an expired certificate, an expensive health endpoint or a backend already unable to carry the transferred load.
The production case is an Azure Front Door Standard or Premium profile with two regional origins. The runbook separates five branches: a bad probe contract, DNS or TLS between Front Door and the origin, an unestablished Private Link path when used, application saturation, and a routing change. It ends with one bounded decision: return the origin to service, keep it drained, fail over explicitly, repair configuration, or restore the last known-good state.
Freeze one origin, route and UTC window
Do not start with weights. Record the profile, endpoint, route, origin group, affected origin and first failure time. A client-side 503 does not prove that a health probe removed the origin. The same status can come from the backend, from having no healthy origin, or from a timeout.
incident:
started_at_utc: 2026-08-31T06:20:00Z
profile: afd-api-prod
endpoint: api-prod
route: api-route
origin_group: api-regions
affected_origin: api-westeurope
standby_origin: api-northeurope
observed:
client_status: 503
origin_health_percentage: <measured-value>
traffic_shift: true-or-false
last_configuration_change: <activity-log-operation>
preserve:
- profile and origin-group configuration
- failed health probe logs by POP
- access logs and X-Azure-Ref samples
- backend health and deployment events
- certificate and Private Link state when applicable
stop_conditions:
- the standby origin approaches its tested capacity
- the candidate origin fails probes from several POPs
- error rate rises after a weight change
- rollback would target an incompatible backend release Also record how many origins are enabled. With one origin, an unhealthy result cannot trigger a failover to a target that does not exist. The probe remains a signal, but it does not create a recovery path.
Read the effective configuration
Probe settings belong to the origin group. Hostname, origin host header, ports, priority, weight and Private Link belong to the origin. The route connects an endpoint to the group. Capture all three objects before changing them.
RG="rg-edge-prod"
PROFILE="afd-api-prod"
ENDPOINT="api-prod"
ROUTE="api-route"
ORIGIN_GROUP="api-regions"
ORIGIN="api-westeurope"
az afd route show \
--resource-group "$RG" \
--profile-name "$PROFILE" \
--endpoint-name "$ENDPOINT" \
--route-name "$ROUTE" \
--output json
az afd origin-group show \
--resource-group "$RG" \
--profile-name "$PROFILE" \
--origin-group-name "$ORIGIN_GROUP" \
--output json
az afd origin show \
--resource-group "$RG" \
--profile-name "$PROFILE" \
--origin-group-name "$ORIGIN_GROUP" \
--origin-name "$ORIGIN" \
--output json
PROFILE_ID=$(az afd profile show \
--resource-group "$RG" \
--profile-name "$PROFILE" \
--query id --output tsv)
az monitor metrics list-definitions \
--resource "$PROFILE_ID" \
--query "[].{metric:name.value,dimensions:dimensions[].value}" \
--output table Check the probe protocol, method, path, interval and sample settings. The path is case-sensitive. Then verify that the origin is enabled, the hostname and origin host header match the backend certificate and virtual host, and the route points to the group under investigation.
Do not relax a probe before explaining its failure. Increasing tolerance reduces detection speed; changing the path changes the contract being tested.
Read failed probes by origin and POP
Azure Front Door health probe logs record failed probes, not a complete stream of successes. Use OriginHealthPercentage for the trend, then use logs to explain the drop. This query targets AzureDiagnostics; adjust field names when the workspace uses resource-specific tables.
let StartTime = datetime(2026-08-31T06:00:00Z);
let EndTime = datetime(2026-08-31T07:00:00Z);
AzureDiagnostics
| where TimeGenerated between (StartTime .. EndTime)
| where Category == "FrontDoorHealthProbeLog"
| extend
Origin = tostring(originName_s),
Result = tostring(result_s),
Status = tostring(httpStatusCode_s),
ProbeUrl = tostring(probeURL_s),
Pop = tostring(POP_s),
OriginIp = tostring(originIP_s),
TotalMs = todouble(totalLatencyMilliseconds_s),
ConnectMs = todouble(connectionLatencyMilliseconds_s),
DnsUs = todouble(DNSLatencyMicroseconds_s)
| where Origin == "api-westeurope.contoso.internal"
| summarize
Failures=count(),
Results=make_set(Result, 10),
Statuses=make_set(Status, 10),
OriginIps=make_set(OriginIp, 10),
P95TotalMs=percentile(TotalMs, 95),
P95ConnectMs=percentile(ConnectMs, 95)
by ProbeUrl, Pop, bin(TimeGenerated, 5m)
| order by TimeGenerated asc, Failures desc Failures limited to a few POPs point toward resolution, a regional path or an intermittent dependency. Simultaneous failures across many POPs strengthen a global hypothesis: certificate, host header, deployment, firewall or backend availability. Repeated 404 or 405 responses usually indicate a bad probe contract. 5xx responses with rising latency require application and dependency evidence.
Separate DNS, TCP, TLS and the application response
A test from an operator workstation does not reproduce Front Door’s path, but it can qualify the origin’s public contract. Test the configured hostname with the expected SNI and host header. Do not rely on the current IP alone.
ORIGIN_HOST="api-we.contoso.net"
ORIGIN_PORT="443"
PROBE_PATH="/health/ready"
dig +short "$ORIGIN_HOST"
openssl s_client \
-connect "$ORIGIN_HOST:$ORIGIN_PORT" \
-servername "$ORIGIN_HOST" \
-verify_return_error \
-brief </dev/null
curl --silent --show-error \
--head \
--connect-timeout 5 \
--max-time 15 \
--resolve "$ORIGIN_HOST:$ORIGIN_PORT:<validated-origin-ip>" \
"https://$ORIGIN_HOST:$ORIGIN_PORT$PROBE_PATH" Interpret each layer separately:
DNSFailureor an unexpected IP set requires checking the origin hostname and its DNS publication;OriginConnectionRefusedor a connection timeout requires listener, filtering and connection-capacity evidence;SSLHandshakeErrororCertificateNameCheckFailedrequires checking chain, expiry, SNI and subject name;- a stable HTTP response rejected by the probe requires correcting the health route or method;
- a slow response or
5xxrequires application health, connection pool and dependency evidence.
The health endpoint should be cheap, deterministic and representative of the ability to serve traffic. It should not run a heavy transaction from every POP, and it should not return healthy when a mandatory dependency is unavailable.
Treat Private Link as one branch, not the default explanation
Azure Front Door Premium can connect to an origin through Private Link. When that path is used, verify that the private endpoint request created for Front Door is approved and the association is established. Health probes follow the same private path as origin traffic.
Private Link does not explain an incident on a public origin, and it does not remove the need to verify hostname, TLS and application behavior. After enabling or changing the association, allow it to converge before drawing a conclusion. Do not reopen public access as the first test: that changes both the security boundary and network path.
Correlate user traffic with backend evidence
Access logs expose OriginName, OriginIP, HttpStatusCode, ErrorInfo, OriginURL and the X-Azure-Ref tracking reference. Use them to determine whether Front Door actually sent requests to the affected backend, whether failures are timeouts, refusals or application errors, and whether the second origin is absorbing the load.
Compare the same UTC window across:
- request and
5xxrates per origin; - Front Door and backend latency;
- CPU, memory, connections, queues and dependency saturation;
- deployment events and Front Door profile changes;
- WAF events only to rule out an upstream block, without confusing WAF behavior with origin health.
An origin that looks healthy in its local portal but receives no Front Door traffic is not end-to-end proof. Conversely, failed probes hidden by cached 200 responses can delay customer impact without repairing the backend.
Decide whether to repair, drain, fail over or roll back
Repair the probe
The path or method no longer matches the health endpoint
The backend serves useful traffic but returns 404 or 405 to probes
The replacement health route is tested without side effects
Repair DNS or TLS
Logs show DNSFailure, SSLHandshakeError or a name mismatch
The expected hostname, SNI and certificate are identified
The fix preserves the origin security contract
Keep the origin drained and repair the backend
Probes fail from several POPs
The backend is saturated or a release regressed
The other origin remains inside its capacity envelope
Fail over or change weights
The second origin is validated and sized
Sessions, data and dependencies support the move
Observation and return criteria are written
Roll back
The incident begins with a Front Door or backend change
The previous state is compatible with current data and certificates
The rollback can be validated with a canary and metrics
Do not change routing
The cause is not localized
The second origin is already near its limit
Probe or access logs are missing Changing priority or weight is not a backend repair. It is a continuity action that needs its own capacity evidence, stop conditions and return path.
Validate recovery without restoring all traffic at once
Re-enable the origin or increase its weight gradually. First require stable probes from multiple POPs, then send a small share of user traffic. Compare errors, latency, connections and business signals with the control origin.
Promote only when OriginHealthPercentage remains within baseline, no new probe failure appears, access logs show the intended origin, and the backend survives a representative traffic cycle. Return to drainage if 5xx, timeouts or queues rise again. Preserve the failed configuration and logs until the incident is closed.
Conclusion
An unhealthy Azure Front Door origin is a chain incident: route, probe, DNS, TCP, TLS, optional Private Link, application and standby capacity. Changing weights too early moves the symptom and can overload the last healthy origin.
A reliable decision comes from one UTC window, failed probes by POP, actual traffic distribution and a gradual return. Repair the contract when it is wrong, drain a genuinely degraded backend, fail over only to a qualified target, and keep rollback open until validation is complete.