Automation
Azure API Management: contain a retry storm before scaling the backend
A production runbook for proving retry amplification in Azure API Management, protecting an overloaded backend, and choosing a bounded policy change, circuit breaker or rollback.
When an API backend starts returning 429, 502, 503 or 504, retries can turn a limited slowdown into an outage. The caller retries, Azure API Management retries, and the backend SDK may retry again. Scaling the backend can briefly absorb the load, but it does not remove the amplification. It can even make the next incident more expensive and harder to explain.
The use case is an order API exposed through Azure API Management. A pricing backend becomes slow during a dependency incident. The APIM policy retries failed calls three times, while clients also retry twice. The runbook goal is to prove where retries occur, reduce pressure without losing evidence, and decide whether to change the retry policy, configure a backend circuit breaker, scale temporarily or restore the previous policy.
Freeze one request path
Start from a single operation and one controlled time window. Do not diagnose the whole API estate from an aggregate 5xx chart.
incident: inc-2026-08-20-012
window_utc:
start: 2026-08-20T06:30:00Z
end: 2026-08-20T07:00:00Z
request:
api: orders-api
operation: GET /orders/{id}/price
correlation_id: retry-check-20260820-01
backend:
apim_backend_id: pricing-backend-prod
expected_timeout_seconds: 8
retry_layers:
client: 2
apim: 3
backend_sdk: unknown
decision_needed:
- reduce_or_remove_apim_retry
- trip_backend_circuit_breaker
- temporary_capacity_change
- rollback_last_policy Record whether the operation is contractually read-only or protected by an idempotency key. A retry on a safe price lookup is not equivalent to retrying an order creation or a payment command.
Calculate the possible amplification
Retry counts multiply across layers. If the client makes one initial attempt plus two retries, and APIM makes one initial backend call plus three retries for every client attempt, one user action can produce up to twelve backend calls before SDK retries are counted.
Maximum backend attempts
client attempts = 1 initial + 2 retries = 3
APIM attempts per client call = 1 initial + 3 retries = 4
possible backend calls = 3 x 4 = 12
This is an upper bound, not an observed count.
Prove actual attempts with correlated backend logs before changing production. This estimate is useful for setting urgency, but it is not evidence by itself. The policy condition may stop early, some clients may not retry, and connection errors may follow a different path.
Read gateway symptoms with KQL
ApiManagementGatewayLogs exposes the client response, backend response, backend duration, total duration and last policy error. It describes one gateway request; it does not automatically prove how many backend attempts occurred inside a retry block.
let Start = datetime(2026-08-20T06:30:00Z);
let End = datetime(2026-08-20T07:00:00Z);
ApiManagementGatewayLogs
| where TimeGenerated between (Start .. End)
| where ApiId == "orders-api" and OperationId == "get-order-price"
| summarize gatewayRequests=count(),
failedGatewayRequests=countif(ResponseCode >= 500 or ResponseCode == 429),
backend429=countif(BackendResponseCode == 429),
backend5xx=countif(BackendResponseCode between (500 .. 599)),
p95BackendMs=percentile(BackendTime, 95),
p95TotalMs=percentile(TotalTime, 95),
errors=make_set(LastErrorReason, 10)
by bin(TimeGenerated, 5m), BackendId
| order by TimeGenerated asc Compare this timeline with backend request logs, dependency metrics and client retry telemetry. A rising BackendTime followed by 429 or 503 supports overload. A stable backend with only gateway policy errors points elsewhere.
Prove each retry at the backend boundary
For a controlled test, add a temporary attempt marker inside the APIM retry block and preserve the incoming correlation identifier. The backend must log both headers. Do this only on the affected operation or a test revision, not globally.
<backend>
<retry
condition="@(context.Response != null && (context.Response.StatusCode == 429 || context.Response.StatusCode == 502 || context.Response.StatusCode == 503 || context.Response.StatusCode == 504))"
count="2"
interval="2"
max-interval="8"
delta="2"
first-fast-retry="false">
<set-variable name="attempt" value="@((context.Variables.ContainsKey("attempt") ? (int)context.Variables["attempt"] : 0) + 1)" />
<set-header name="X-APIM-Attempt" exists-action="override">
<value>@(((int)context.Variables["attempt"]).ToString())</value>
</set-header>
<set-header name="X-Correlation-Id" exists-action="skip">
<value>@(context.RequestId.ToString())</value>
</set-header>
<forward-request buffer-request-body="true" />
</retry>
</backend> The retry policy executes its child policies once before evaluating whether another attempt is needed. Therefore count="2" permits two retries after the initial execution. Do not infer that count is the total number of backend calls.
Separate retryable failures from unsafe retries
An operational retry policy needs an explicit failure contract.
Read-only operation, short 502/503/504
Allow a small retry count with backoff and a total latency budget
429 with Retry-After
Respect backend recovery time
Prefer circuit breaking to repeated immediate calls
POST or side-effecting operation without idempotency proof
Do not add an automatic retry
Reconcile backend state before any replay
Timeout with unknown backend state
Treat the outcome as ambiguous
Query the business record before retrying
Policy error before the backend call
Fix the policy or identity path
Backend scaling will not help The key distinction is not only HTTP status. It is whether the previous attempt may have changed state and whether the next attempt remains inside a documented latency and capacity budget.
Contain pressure without hiding the incident
The first production action should reduce amplification while keeping the diagnostic path alive. For a read-only operation, lower the retry count, disable the fast first retry and use backoff. For an unsafe operation, remove the retry and return an explicit failure while reconciliation runs.
A backend circuit breaker is appropriate when repeated failures should stop calls for a recovery period. In APIM, a tripped backend circuit returns 503 until it resets. Its decision is approximate across distributed gateway instances, so it is a protection mechanism, not a precise global counter. It should be combined with client guidance, observability and a tested recovery path.
action:
operation: get-order-price
mode: bounded_retry
retry_count: 1
first_fast_retry: false
total_latency_budget_ms: 12000
backend_protection:
circuit_breaker: evaluate
failure_codes: [429, 500-599]
accept_retry_after: true
still_enabled:
- gateway diagnostics
- backend request logs
- synthetic read-only probe
blocked:
- retries on order creation
- global APIM policy change
- capacity increase without an expiry time Do not deploy a broad rate limit, retry or circuit breaker from an incident sample alone. Scope it to the backend and operations whose failure contract is understood.
Decide change, scale or rollback
Use three signals together: retry amplification, backend saturation and business impact.
Change the retry policy when
Correlated backend logs prove repeated attempts
The operation is retry-safe
Backoff and latency budget are defined
Configure or adjust the circuit breaker when
The backend needs a recovery window
Failure conditions and trip duration are testable
Clients handle the resulting 503 and Retry-After path
Scale temporarily when
Demand is legitimate after retry amplification is contained
Capacity has an owner, limit and expiry time
Rollback when
The storm started after a policy deployment
Guardrail probes fail or side effects are duplicated
The previous policy is known and deployable After the change, replay one positive probe and one failure probe. Confirm the healthy call succeeds, the failure does not exceed the approved attempt count, backend pressure falls, correlation remains visible and unsafe operations are not retried. If any guardrail fails, restore the previous policy and keep the backend protection decision open.
Conclusion
A retry storm is an amplification incident, not only a capacity incident. The useful evidence is one request path, the retry count at every layer, correlated gateway and backend logs, the operation’s idempotency contract and a bounded latency budget.
The final decision then becomes defensible: reduce retries when APIM amplifies the failure, use a circuit breaker when the backend needs recovery time, scale only after amplification is contained, or roll back the last policy when it introduced the storm. The objective is not to make every call eventually succeed. It is to keep failure controlled, observable and reversible.