Automation

Azure Logic Apps: contain connector throttling before adding retries

A production runbook for separating Logic Apps limits, connector throttling and downstream saturation, then reducing concurrency, validating a canary or rolling back without amplifying 429 responses.

27 Aug 2026 azurelogic-appsconnectorsthrottling429retryconcurrencyobservabilitykqlautomationrunbookrollbackproduction

An Azure Logic Apps workflow synchronizes orders with a SaaS API. After a burst of messages, some actions return 429 Too Many Requests, runs remain active for much longer, and the backlog grows. The immediate response is often to add retries or raise concurrency to catch up. Both changes can multiply calls to a dependency that is already saturated.

The runbook must first locate the limit: the Logic Apps resource, the connector connection, or the destination system. It must then contain pressure, preserve ordering and idempotency, and decide between lower concurrency, a bounded retry, queue buffering, additional capacity, or rollback.

Freeze one representative window and path

Do not start from a total count of failed runs. Select one workflow, action, connection, and UTC window where the symptom can be reproduced.

yaml logic-app-throttling-contract.yml
incident: inc-2026-08-27-logicapps-429
window_utc:
start: 2026-08-27T05:40:00Z
end: 2026-08-27T06:10:00Z
workflow:
resource: la-orders-prod
name: export-orders
trigger: service-bus
action: create_order_in_saas
connection: saas-orders-prod
symptoms:
- action_returns_429
- runs_stay_running
- source_backlog_grows
recent_change:
- trigger_concurrency_increased
decision:
- contain_parallelism
- keep_bounded_retry
- buffer_or_scale
- rollback_last_workflow_revision

Keep the run ID, business identifier, action name, attempt count, start and end times, and any Retry-After header. Without that scope, an average can easily hide one constrained connector or a single slow operation.

Separate the three throttling layers

A 429 does not automatically identify the SaaS service as the source. The diagnosis must distinguish three layers.

text throttling-layers.txt
Logic Apps resource
Several workflows or runs consume execution capacity
Signals: throttled events, waiting runs, broad increase in 4xx responses

Connector or API connection
One operation exceeds a connector or connection limit
Signals: 429 responses concentrated on one action and connection

Destination system
The API, database or SaaS rejects calls it cannot absorb
Signals: Retry-After, destination quotas, application logs and latency

Required proof
Match one Logic Apps run to one destination request
Compare time, correlation ID, action, connection and attempt

For a Consumption workflow, inspect throttled action and trigger events, then verify the connection used by the action. For a Standard workflow, correlate HTTP 4xx metrics, runtime telemetry, and destination logs. The hosting model changes the available signals, but not the requirement to prove which layer rejected the request.

Read attempts before changing the policy

Run history shows whether an action was retried. Diagnostics sent to Log Analytics extend that view across multiple executions.

kusto 01-logic-app-retry-history.kql
let Start = datetime(2026-08-27T05:40:00Z);
let End = datetime(2026-08-27T06:10:00Z);
LogicAppWorkflowRuntime
| where TimeGenerated between (Start .. End)
| where WorkflowName == "export-orders"
| extend HasRetry = isnotempty(RetryHistory)
| summarize operations=count(),
          operationsWithRetry=countif(HasRetry),
          failures=countif(Status =~ "Failed"),
          sampleErrors=make_set(Error, 5)
by bin(TimeGenerated, 5m), OperationName
| order by TimeGenerated asc

Open a few representative runs and inspect the action inputs, outputs, and attempts while protecting sensitive data. An increase in operationsWithRetry alongside backlog growth indicates possible amplification; it does not yet prove destination saturation. Correlate the window with dependency logs and quotas.

Calculate concurrency amplification

Actual pressure comes from the combination of parallel runs, For each loops, Split On, and retries. An upper-bound estimate is enough to expose a dangerous change.

text logic-app-amplification.txt
Example upper bound
20 concurrent runs
10 parallel For each iterations per run
1 initial call + 4 retries per action

possible calls = 20 x 10 x 5 = 1,000

This is not an observed volume.
It tests whether configuration can overwhelm the destination.
Logs must prove the real number of attempts.

Include upstream client retries and downstream SDK retries when they exist. Several layers retrying independently can turn a limited slowdown into a request wave even when each local policy looks reasonable.

Contain pressure without losing the message

The first production action should reduce new calls while keeping data recoverable. Lower trigger or loop concurrency within a known scope. Disable Split On temporarily only if the workflow can process the array without losing its recovery unit. Do not delete source messages to make a chart fall.

json bounded-concurrency-and-retry.json
{
"triggers": {
  "When_message_received": {
    "runtimeConfiguration": {
      "concurrency": {
        "runs": 4
      }
    }
  }
},
"actions": {
  "Create_order_in_SaaS": {
    "inputs": {
      "retryPolicy": {
        "type": "exponential",
        "count": 2,
        "interval": "PT10S",
        "minimumInterval": "PT10S",
        "maximumInterval": "PT1M"
      }
    }
  }
}
}

This fragment illustrates intent, not a definition to paste without checking the exact trigger and action type. Some operations expose a retry policy and others do not. Version the complete workflow, review the diff, and keep the previous definition ready before deployment.

When the source is a queue, slowing the consumer is often safer than forcing throughput. The backlog becomes an observable buffer. Define its maximum acceptable age, capacity, retention window, and recovery owner.

Choose retries from the action contract

A retry is acceptable only when the previous effect is known or can be reconciled.

text retry-safety-matrix.txt
Idempotent GET with 429 and Retry-After
Use a bounded retry compatible with the workflow time budget

POST protected by an idempotency key
Retry only after verifying the key at the destination

POST without idempotency, timeout or ambiguous response
Do not retry automatically
Reconcile destination state before recovery

429 with no proven source
Reduce concurrency and complete the diagnosis
Do not multiply connections to bypass an unknown quota

Local error before the destination call
Fix the workflow or connector
More downstream capacity will not help

Honoring Retry-After is not enough if hundreds of runs wake at the same instant. Add jitter where the mechanism supports it, bound the attempt count, and define a total time limit. Beyond that limit, the message must remain available for controlled recovery or move to an operable failure path.

Validate with a canary, then resume in stages

Deploy containment to a test workflow or revision and inject a small identifiable batch. The canary must prove both the healthy path and behavior under 429.

yaml logic-app-canary.yml
canary:
messages: 10
correlation_prefix: logicapps-429-canary
acceptance:
- no_duplicate_business_effect
- retry_count_at_or_below_approved_limit
- target_429_rate_decreases
- backlog_age_stabilizes
- end_to_end_correlation_remains_visible
resume_steps:
- concurrency_2
- concurrency_4
- concurrency_8_only_if_target_stays_healthy
stop_conditions:
- duplicate_detected
- retry_budget_exceeded
- target_latency_rises_again
- backlog_cannot_be_drained_before_retention_limit

A green run is not sufficient validation. Check the business effect at the destination, the absence of duplicates, backlog age, and the real call count per identifier.

Decide correction, capacity, or rollback

Keep lower concurrency when it holds throughput within proven destination capacity. Change the retry policy only when the action is safe, total delay is bounded, and attempts are visible. Add a queue or another decoupling boundary when bursts are legitimate but incompatible with downstream throughput. Add capacity only after artificial amplification is removed.

Roll back the latest definition when the incident follows a concurrency increase, Split On activation, or retry change and the previous version is known. Rollback must restore the definition, verify connections, run a canary, and drain the backlog in stages. It must not delete messages or hide ambiguous runs.

Conclusion

A 429 in Azure Logic Apps is a distributed capacity decision across runtime, connector, and destination. Adding retries before identifying that boundary risks consuming more of the resource that is already asking callers to slow down.

The useful runbook correlates one run with one destination request, measures attempts, calculates possible concurrency, contains throughput, and proves idempotency. The final decision then becomes operable: lower concurrency, retain a bounded retry, introduce buffering, add proven capacity, or roll back the workflow definition that triggered amplification.