AI
Azure API Management AI gateway: validate a backend circuit breaker before production failover
A production runbook for qualifying circuit breakers, Retry-After, Microsoft Foundry backend pools, canary traffic, observability and rollback before enabling AI API failover.
An AI API exposed through Azure API Management calls a primary model deployment and keeps a second Microsoft Foundry endpoint for recovery. During a demand spike, the primary starts returning 429 responses and intermittent 5xx errors. Placing both endpoints in a pool and enabling a circuit breaker looks like enough to preserve service.
It can just move the incident. A sensitive threshold opens the circuit on a short burst. A permissive threshold lets retries amplify saturation. A smaller secondary receives the entire workload, or a 401 identity failure is mistaken for a transient outage. This runbook builds a failover decision from the traffic contract, deployed configuration, a bounded canary and evidence from APIM and each backend.
Freeze the contract before changing the pool
Describe the complete production unit: API, operation, backends, model deployments, quotas, identity and client behavior. A circuit breaker protects a backend. It cannot repair a mismatched token budget, an invalid identity or an unbounded retry policy.
gateway: apim-ai-prod
api: assistant-operations-v1
operation: POST /responses
backends:
primary: foundry-model-weu
secondary: foundry-model-neu
failure_signals:
transient:
- 429-with-retry-after
- 500-to-599
configuration:
- 401
- 403
client:
- 400
- 404
canary_share_percent: 5
max_end_to_end_latency_ms: 12000
max_attempts_across_all_layers: 2
promote_requires:
- secondary-model-and-deployment-validated
- identity-proven-on-both-backends
- circuit-opens-and-recovers-as-designed
- no-retry-amplification
rollback_on:
- secondary-saturation
- semantic-regression
- unexplained-401-or-403
- missing-correlation-evidence A Private Endpoint may secure the path to a Foundry endpoint, but it does not determine backend health. Validate DNS and routing separately; the circuit breaker must not hide a broken resolution or network path.
Read the deployed backend resources
Inventory the APIM backend resources through the management API before making a change. Preserve the response and its ETag with the change record. Together they are the rollback baseline.
SUBSCRIPTION_ID="00000000-0000-0000-0000-000000000000"
RG="rg-ai-platform-prod"
APIM="apim-ai-prod"
API_VERSION="2024-05-01"
BASE="https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.ApiManagement/service/$APIM"
az rest --method get --url "$BASE/backends?api-version=$API_VERSION" --output json > apim-backends-before.json
jq '.value[] | {
name,
type: .properties.type,
url: .properties.url,
pool: .properties.pool,
circuitBreaker: .properties.circuitBreaker
}' apim-backends-before.json Confirm that the pool references the intended ARM IDs rather than similarly named resources. Compare priority, weight, region, model and deployment. A higher-priority recovery backend must be able to accept the expected traffic when the primary leaves the pool.
Select the failures that should trip the circuit
For a model API, 429 and 5xx responses can describe transient saturation or unavailability. 401 and 403 usually identify an audience, role, credential or policy problem. Failing over on them can conceal identity drift without fixing it. Client-generated 4xx responses should not remove a healthy backend from the pool.
Tie the threshold to volume. Three failures out of five requests carry a different signal from three failures out of fifty thousand. Use a count for a predictable flow or a percentage when volume varies, then exercise it with a representative sample. acceptRetryAfter lets the backend influence the open duration through Retry-After. Enable it only after proving that the header is present, consistent and bounded in your test cases.
{
"properties": {
"title": "Foundry model West Europe",
"description": "Primary production model endpoint",
"protocol": "http",
"url": "https://foundry-model-weu.example.azure.com",
"circuitBreaker": {
"rules": [
{
"name": "transient-model-failures",
"failureCondition": {
"count": 5,
"interval": "PT30S",
"statusCodeRanges": [
{ "min": 429, "max": 429 },
{ "min": 500, "max": 599 }
]
},
"tripDuration": "PT30S",
"acceptRetryAfter": true
}
]
}
}
} This is a test starting point, not a universal threshold. APIM backend circuit breakers are not supported in the Consumption tier. Confirm the service tier and supported management API version before rollout.
Prove that the secondary is a real recovery backend
A declared secondary endpoint is not recovery capacity yet. Run the same canary prompt, with the same identity and essential parameters, directly against each backend and then through APIM. At minimum, verify the expected model and version, token limits, safety filters, latency, region, network access and logging.
Make the canary as deterministic as the model allows and free of business side effects. It can request a short JSON response whose schema is validated mechanically. For an agent, disable state-changing tools during this test. The goal is to qualify the model path, not execute an operational runbook.
{
"input": "Return only a JSON object with status=ready and contract=ai-gateway-v1.",
"max_output_tokens": 40,
"metadata": {
"test": "apim-circuit-breaker-canary",
"change": "chg-2026-09-14-01"
}
} A 200 is not enough. Validate the schema, expected deployment marker and correlated trace. If the secondary responds but applies another filter, model or output contract, failover is technically available and functionally invalid.
Exercise opening, diversion and recovery
Deploy the rule to one backend and a canary API first. Generate a bounded sequence of simulated responses or use a controlled test backend. Do not create real model saturation. The test must expose four states: normal traffic, threshold reached, backend excluded, and traffic restored after the open duration.
Case A - baseline
Primary returns 200
Expected: primary selected, one backend call, schema valid
Case B - below threshold
Primary returns fewer failures than configured count
Expected: circuit remains closed, no uncontrolled retry burst
Case C - threshold reached
Primary returns qualified 429 or 5xx responses
Expected: circuit opens, pool selects eligible secondary, clients see bounded behavior
Case D - configuration error
Primary returns 401 or 403
Expected: no failover caused by circuit rule, identity incident remains visible
Case E - recovery
Trip duration expires and primary is healthy
Expected: traffic resumes according to pool priority and weight, without oscillation For every case, capture UTC time, correlation ID, selected backend, APIM status, backend status, latency and total attempts. When the circuit is open and no eligible backend remains, APIM can return 503. That outcome belongs in the client contract; it should not be discovered during an incident.
Correlate APIM, model and retries
Table names vary with the configured diagnostic destination and mode. Start with the schema actually ingested, then build a view that separates gateway and backend responses. The query below illustrates the required result rather than promising a universal table name.
let Start = datetime(2026-09-14T14:00:00Z);
let End = datetime(2026-09-14T14:30:00Z);
ApiManagementGatewayLogs
| where TimeGenerated between (Start .. End)
| where ApiId == "assistant-operations-v1"
| extend CorrelationId = tostring(CorrelationId)
| project TimeGenerated,
CorrelationId,
BackendId,
ResponseCode,
BackendResponseCode,
TotalTime,
BackendTime,
LastErrorReason,
LastErrorMessage
| order by TimeGenerated asc Group results by correlation and a short time window. More backend calls per client request reveal retry amplification. An APIM 503 without a backend attempt during the open window can be expected. A backend 401 must remain visible as an identity incident. A secondary 200 with latency or token consumption outside budget is not successful recovery.
Decide rollout, adjustment or rollback
deploy:
when:
- canary-proves-primary-and-secondary-contracts
- only-transient-failures-trip-the-circuit
- secondary-capacity-is-bounded-and-observed
- retry-count-remains-within-contract
- recovery-does-not-oscillate
adjust:
when:
- threshold-is-too-sensitive-for-observed-volume
- retry-after-exceeds-operational-budget
- pool-weight-overloads-secondary
action:
- change-one-parameter
- replay-the-same-test-matrix
hold:
when:
- backend-identity-or-model-version-is-ambiguous
- diagnostics-cannot-identify-selected-backend
- secondary-has-no-proven-capacity
rollback:
action:
- restore-backend-resources-from-reviewed-iac
- restore-previous-pool-priorities-and-weights
- remove-only-the-new-circuit-rule
- replay-baseline-and-identity-negative-test
- confirm-client-visible-errors-return-to-baseline Do not change threshold, duration, pool weight and retries in one step. Without an isolated variable, the canary proves nothing and rollback becomes approximate. Keep the configuration in IaC with a contract test and an owner for threshold changes.
Conclusion
An APIM circuit breaker does not automatically turn two Microsoft Foundry endpoints into a resilient service. Teams still need to qualify transient failures, protect the secondary, bound retries and prove that recovery preserves identity, model, safety and output contract.
Roll out when the canary demonstrates opening, traffic diversion and recovery without amplification or oscillation. Hold when the selected backend or runtime identity remains ambiguous. If the secondary saturates or semantics change, restore the previous backend resource and pool weights, then replay the same matrix before closing the change.