Automation

Azure Logic Apps: diagnose duplicate runs before disabling retries

A production runbook for separating redelivery, splitOn, concurrency and ambiguous retries, then enforcing idempotency without removing workflow resilience.

24 Sept 2026 azurelogic-appsautomationidempotencyretriesconcurrencysplitonobservabilityincidentrunbookrollbackproduction

An Azure Logic Apps workflow creates two downstream orders for one business event. Both runs succeeded, their start times are close, and the source system shows a single user action. Disabling retries or setting trigger concurrency to one looks like an immediate fix. Either change might reduce the symptom, but neither proves where the duplicate originated nor prevents another duplicate after the next ambiguous timeout.

The running case is a workflow that receives an order event, enriches its payload, and calls an internal API. Azure Logic Apps uses at-least-once message delivery, so the integration boundary must preserve business integrity when a message is occasionally delivered again. This runbook links every downstream effect to an event, trigger occurrence, workflow run, and action attempt before choosing a producer fix, throughput adjustment, idempotency barrier, or rollback.

Freeze one duplicate before changing the workflow

Start with a pair of effects that are proven duplicates. Preserve the business identifier, UTC times, both run IDs, trigger history, workflow revision, and identifiers returned by the target. Two identical rows on a dashboard are not enough: they might represent two valid events, two views of the same effect, or two separate writes.

yaml logic-app-duplicate-scope.yml
business_event:
event_id: order-84721
entity_id: customer-order-392
produced_at_utc: 2026-09-24T08:14:02Z

workflow:
resource: la-orders-prod
name: process-order
definition_version: git-sha-or-release-id
runs:
  - run-a
  - run-b

downstream_effects:
- request_id: api-request-1901
  result_id: fulfillment-551
- request_id: api-request-1908
  result_id: fulfillment-552

temporary_guards:
- do not resubmit either run
- do not purge run history
- do not disable retries globally
- stop manual replay for this event

Export the relevant inputs and outputs while redacting secrets and personal data. Record recent changes to the trigger, splitOn, concurrency, retry policy, loops, and target API. The evidence must outlive run-history retention.

Build an identifier chain

A duplicate becomes explainable when four levels are connected: the business event emitted by the source, the trigger occurrence, the workflow run, and the attempt sent to the downstream system. Start with an existing stable domain identifier and propagate it unchanged.

text duplicate-correlation-matrix.txt
Event ID      Trigger occurrence   Workflow run   Action attempt   Downstream result
order-84721   trigger-301          run-a          attempt-1        fulfillment-551
order-84721   trigger-302          run-b          attempt-1        fulfillment-552

Questions
- Did the source emit the same event ID twice?
- Did one trigger create multiple runs through splitOn?
- Did one run repeat an action after a timeout?
- Did the target process the same business key twice?

Do not substitute the workflow run name for the business identifier. A redelivery creates a new technical run while representing the same intent. Conversely, two separate business events must not be collapsed merely because their visible payloads look similar.

Separate four duplicate paths

The first path is upstream: a webhook sent twice, an event republished after a lost acknowledgement, a poller rereading an overlapping window, or a producer that does not retain its identifier. The two runs then have separate trigger occurrences but the same business key.

The second path is debatching. With splitOn, every array item starts a separate workflow run. Inspect the exact payload and split expression. If two array elements carry the same key, the platform deliberately creates two runs; reducing concurrency does not merge them.

The third path is overlapping execution. Two valid occurrences can run concurrently, read the same unchanged state, and both write. Trigger concurrency set to one may serialize this workflow, but it does not protect against later redelivery or a second producer instance.

The fourth path is an ambiguous retry. An HTTP action times out from the workflow’s point of view after the target has committed the write. A retry then sends the same intent again. Disabling retries removes that second attempt, but it also turns transient failures and 429 responses into loss or manual recovery. The durable control belongs at the boundary that creates the effect.

Read retries as attempts, not runs

Open both runs and find the first divergence. If each run contains one successful attempt, investigate the trigger and producer. If one run contains multiple attempts for an action after a timeout, 408, 429, or 5xx response, correlate every attempt with target-side logs. A failed orchestrator status does not prove that the server produced no effect.

Capture the effective retry policy, including connector defaults. Compare the client timeout with target processing time. An API that commits and answers late creates an ambiguity window; only increasing the timeout moves that window without making the operation idempotent.

Enforce idempotency at the business boundary

The target should recognize an intent it has already processed and return the original result without creating the effect again. The key must remain stable for the same business operation and change for a new deliberate operation. A useful shape is orderId + operation + version, not a timestamp or run ID.

json idempotent-request-contract.json
{
"eventId": "order-84721",
"operation": "create-fulfillment",
"entityVersion": 3,
"idempotencyKey": "order-84721:create-fulfillment:v3",
"correlation": {
  "workflowRun": "run-a",
  "action": "Create_fulfillment"
}
}

The barrier must be atomic. A separate “does this key exist?” read followed by a write is still vulnerable to concurrent runs. Prefer conditional creation, a unique constraint, or a transaction that stores the key and effect together. Retain the result, status, and a lifetime that covers the maximum redelivery window.

If the target cannot change immediately, place a shared idempotency barrier before the action. Treat it as a compensating control: storage failure, premature expiry, a stuck in_progress state, and crash recovery all need defined behavior.

text idempotency-state-machine.txt
absent -> in_progress -> completed
                  -> failed_retryable
                  -> failed_final

On duplicate key
completed        return stored result; do not write again
in_progress      wait or return accepted; do not start a second write
failed_retryable retry with the same key and bounded ownership
failed_final     stop and require an explicit decision

Canary without creating a third effect

First deploy key propagation and observation without removing retries. For a controlled test order, deliver the same event twice and, when the test environment permits it, delay the response after a successful write. The expected result is one downstream effect, two correlated traces, and a stable response to the second request.

Then run the negative test: two valid operations on the same entity with different versions or intents must both proceed. An overly broad key eliminates duplicates by suppressing legitimate changes as well.

During the canary, count events, trigger occurrences, runs, attempts, and downstream effects by key. These counts do not have to match: splitOn and retries can increase intermediate levels. The invariant is one business effect per idempotent intent.

Decide, validate, or roll back

Fix the source when it emits duplicate events without a valid reason. Correct splitOn when the array or expression produces unexpected work units. Reduce concurrency when the target is saturated, but do not present serialization as a uniqueness guarantee. Keep retries for transient failures when the downstream boundary can safely absorb repetition.

Roll back if the new key merges distinct operations, if the idempotency store becomes a failure point, or if replayed responses do not match the original result. Restore the previous workflow revision, block resubmissions, and keep a bounded manual procedure while correcting the design. Do not delete idempotency records already created; they still protect against in-flight events.

Final validation needs three proofs: replaying the same event creates no new effect, a new legitimate intent is still accepted, and a timeout after commit converges on the existing result. Record the key retention period, its owner, and the recovery procedure for a stuck state.

Conclusion

A duplicate Logic Apps run is not automatically a broken retry. It can originate in the producer, splitOn, overlapping runs, or a response lost after a real effect. The identifier chain reveals the first level that duplicates before several settings are changed at once.

The runbook ends with an observable decision: repair the producer, correct debatching, bound concurrency, make the target idempotent, or roll back the barrier. Resilience remains valuable when repeating an attempt no longer means repeating the business effect.