AI

Azure API Management AI gateway: diagnose token-limit 429s before raising the quota

A production runbook for separating APIM counters, policy scope, token estimation, backend throttling and client retries before expanding an AI token limit.

26 Sept 2026 azureapi-managementapimai-gatewayllm-token-limitmicrosoft-foundryrate-limitingtokensobservabilityguardrailsrunbookrollbackproduction

An AI API exposed through Azure API Management starts returning 429 responses after an llm-token-limit policy rollout. The Microsoft Foundry backend still has capacity, yet some tenants are blocked while others continue normally. Raising tokens-per-minute looks like the fastest recovery.

That change can hide a broken isolation key, an empty value shared by every consumer, duplicate policy inheritance, gateway-specific counters or a retry loop that turns healthy throttling into an incident. The running case is a multi-tenant POST /responses API. This runbook attributes the 429, reconstructs the effective budget, canaries a correction and ends with a promote, hold or rollback decision.

Freeze one rejected request and the traffic contract

Preserve one complete example before editing the policy: UTC timestamp, gateway, region, API, operation, consumer identity, requested model, status, Retry-After, correlation ID and approximate prompt size. Do not log raw prompts, secrets or personal data just to make diagnosis easier.

Then write down the intended budget. A per-tenant limit, a per-application limit and a global safety ceiling protect different things.

yaml ai-token-limit-contract.yml
gateway: apim-ai-prod
api: assistant-v1
operation: POST /responses
consumer_key_source: validated tenant_id claim
rate_budget:
tokens_per_minute: 30000
quota_budget:
tokens: 500000
period: Daily
client_contract:
honors_retry_after: true
max_attempts: 2
jitter: true
promotion_requires:
- one-counter-per-tenant
- gateway-429-distinguished-from-backend-429
- no-empty-counter-key
- canary-within-budget
rollback:
restore: apim-policy-before.xml
verify: previous-counter-behavior-and-backend-traffic

These numbers describe the example contract, not a universal recommendation. Derive them from backend capacity, accepted cost, consumer count and observed request sizes.

Prove which layer produced the 429

A client-visible 429 can come from the APIM policy, the model backend or another limiter in the path. Correlate configured response headers, gateway traces and backend logs using the same request identifier.

Give policy headers explicit names so operators can observe the decision without exposing the internal counter key.

xml llm-token-limit-policy.xml
<llm-token-limit
counter-key="@(context.Principal?.Claims.GetValueOrDefault(&quot;tenant_id&quot;, &quot;&quot;))"
tokens-per-minute="30000"
token-quota="500000"
token-quota-period="Daily"
estimate-prompt-tokens="true"
retry-after-header-name="x-ai-retry-after"
remaining-tokens-header-name="x-ai-tokens-remaining"
remaining-quota-tokens-header-name="x-ai-quota-remaining" />

When APIM rejects a request before forwarding it, the backend should have no matching call. If the backend returns its own 429, the policy may be working correctly: inspect deployment capacity, provider quota, backend pools and circuit breakers. Do not merge both sources into one chart called “throttling.”

Audit effective scope and policy order

Policies can be inherited across global, workspace, product, API and operation scopes. Reconstruct the effective policy, including base elements, before assuming that only one limiter ran. Two valid limits can coexist, but each needs a distinct purpose, key and budget.

Look for five common failure modes:

  1. a global limit and an API limit using the same key without a scope convention;
  2. inconsistent tokens-per-minute values for the same key on v2 tiers;
  3. a portal-created policy duplicating the version already deployed from IaC;
  4. a limiter executed before the identity validation that supplies counter-key;
  5. an expression returning an empty string when the expected claim is absent.

An empty key is an isolation incident because unrelated tenants feed the same counter. Reject the request before the limiter when the validated identifier is missing. Do not silently fall back to a shared value such as unknown.

xml reject-missing-tenant-before-counter.xml
<choose>
<when condition="@(string.IsNullOrWhiteSpace(context.Principal?.Claims.GetValueOrDefault(&quot;tenant_id&quot;, &quot;&quot;)))">
  <return-response>
    <set-status code="401" reason="Validated tenant identity required" />
  </return-response>
</when>
</choose>

Recalculate the budget that is actually consumed

A token limit is not a request limit. Two calls can count as one request each while consuming radically different token volumes. The policy uses usage returned by the model and can estimate prompt tokens before forwarding, preventing calls that are already over budget. Remaining-token values are estimates, however; they are not an exact financial ledger.

Group evidence by tenant, model, operation and gateway with bounded cardinality. Avoid request IDs, conversation IDs and individual users as metric dimensions. They make the series expensive and eventually hide new values.

For streaming responses, verify that usage is actually returned. An interrupted stream or a provider that omits usage makes measurement incomplete; missing telemetry does not mean zero consumption. Exercise streaming and non-streaming separately under controlled load.

xml emit-bounded-token-metrics.xml
<llm-emit-token-metric namespace="NaxayaAIGateway">
<dimension name="API ID" />
<dimension name="Operation ID" />
<dimension name="Gateway ID" />
<dimension name="ConsumerTier" value="@(context.Request.Headers.GetValueOrDefault(&quot;x-consumer-tier&quot;, &quot;unknown&quot;))" />
</llm-emit-token-metric>

The tier header must come from a trusted identity or controlled mapping, not a free-form client value.

Account for gateways and regions

Counters are tracked independently at each gateway. A multi-region instance, workspace gateway and self-hosted gateway do not form one global counter. The same tenant can therefore be admitted in one region and limited in another based on local traffic.

Break down 429 responses, tokens and requests by gateway before raising a limit. Check routing changes as well: failover, lost affinity or regional recovery can suddenly concentrate traffic on a different counter. If the business requirement is one strict global budget, document that a local policy is insufficient on its own and place the aggregate control in an appropriate layer.

Verify that Retry-After does not create a retry storm

Correct throttling can become an outage when the SDK, gateway and agent each retry. Measure original calls, retries and tokens per business action. Clients should honor Retry-After, add jitter and cap total attempts across the whole chain.

Do not automatically retry a non-idempotent request that can invoke a tool or perform a write. For an agent workflow, retain the business action ID and tool-call state so a retry cannot replay an operation that was already accepted.

text retry-decision.txt
Gateway 429 with Retry-After
wait once with jitter
keep the same business action id
stop when the end-to-end retry budget is exhausted

Backend 429
inspect deployment capacity and gateway routing
do not raise the APIM token budget as a substitute

Missing tenant identity
fail authentication
never fall back to a shared counter

Write-capable agent action
verify prior tool-call state before any retry

Canary a correction without resetting the diagnosis

Create a bounded canary route or product with an explicitly distinct counter key, for example by suffixing the policy version. Send one test tenant and exercise three profiles: steady small prompts, a controlled burst and a request intentionally over budget.

The canary must prove that:

  1. tenant A does not consume tenant B’s budget;
  2. 429 appears at the intended boundary with a usable Retry-After;
  3. APIM does not call the backend when it blocks early;
  4. admitted traffic produces coherent metrics;
  5. the client stops within its retry budget;
  6. agent actions with side effects remain protected from replay.

Do not deploy a fresh counter-key to all production traffic merely to clear the 429 responses. That effectively resets counters and destroys evidence. Keep the new key inside the bounded canary.

Decide: correct, expand or roll back

Correct the policy when tenants share a key, scope is duplicated, identity is not validated before counting, or one key carries inconsistent values. Available backend capacity does not justify broken isolation.

Raise the budget only after key and scope are proven, telemetry shows legitimate demand, the backend can absorb it and the cost owner accepts the change. Roll out in stages with stop conditions for backend 429, latency, errors and consumption.

Roll back when the canary mixes consumers, telemetry becomes incomplete, retries amplify or the correction merely moves rejection to the backend. Restore the exported policy, verify effective policy at every scope, then replay one positive and one limited test. Rollback is complete only when counter behavior, backend traffic and client headers match the last known contract.

Conclusion

A 429 after adding llm-token-limit is not automatically evidence that the quota is too low. It can expose a shared key, duplicated inheritance, gateway-local behavior, incomplete token accounting or retries amplifying demand.

The incident becomes operable when the team can name which layer rejected the call, which counter was charged and why. Repair isolation and evidence first; expand the budget only with proven backend capacity and accepted cost; roll back as soon as the canary can no longer attribute consumption and rejection precisely.