Automation
Azure Logic Apps: diagnose duplicate runs before disabling retries
A production runbook for separating redelivery, splitOn, concurrency and ambiguous retries, then enforcing idempotency without removing workflow resilience.
An Azure Logic Apps workflow creates two downstream orders for one business event. Both runs succeeded, their start times are close, and the source system shows a single user action. Disabling retries or setting trigger concurrency to one looks like an immediate fix. Either change might reduce the symptom, but neither proves where the duplicate originated nor prevents another duplicate after the next ambiguous timeout.
The running case is a workflow that receives an order event, enriches its payload, and calls an internal API. Azure Logic Apps uses at-least-once message delivery, so the integration boundary must preserve business integrity when a message is occasionally delivered again. This runbook links every downstream effect to an event, trigger occurrence, workflow run, and action attempt before choosing a producer fix, throughput adjustment, idempotency barrier, or rollback.
Freeze one duplicate before changing the workflow
Start with a pair of effects that are proven duplicates. Preserve the business identifier, UTC times, both run IDs, trigger history, workflow revision, and identifiers returned by the target. Two identical rows on a dashboard are not enough: they might represent two valid events, two views of the same effect, or two separate writes.
business_event:
event_id: order-84721
entity_id: customer-order-392
produced_at_utc: 2026-09-24T08:14:02Z
workflow:
resource: la-orders-prod
name: process-order
definition_version: git-sha-or-release-id
runs:
- run-a
- run-b
downstream_effects:
- request_id: api-request-1901
result_id: fulfillment-551
- request_id: api-request-1908
result_id: fulfillment-552
temporary_guards:
- do not resubmit either run
- do not purge run history
- do not disable retries globally
- stop manual replay for this event Export the relevant inputs and outputs while redacting secrets and personal data. Record recent changes to the trigger, splitOn, concurrency, retry policy, loops, and target API. The evidence must outlive run-history retention.
Build an identifier chain
A duplicate becomes explainable when four levels are connected: the business event emitted by the source, the trigger occurrence, the workflow run, and the attempt sent to the downstream system. Start with an existing stable domain identifier and propagate it unchanged.
Event ID Trigger occurrence Workflow run Action attempt Downstream result
order-84721 trigger-301 run-a attempt-1 fulfillment-551
order-84721 trigger-302 run-b attempt-1 fulfillment-552
Questions
- Did the source emit the same event ID twice?
- Did one trigger create multiple runs through splitOn?
- Did one run repeat an action after a timeout?
- Did the target process the same business key twice? Do not substitute the workflow run name for the business identifier. A redelivery creates a new technical run while representing the same intent. Conversely, two separate business events must not be collapsed merely because their visible payloads look similar.
Separate four duplicate paths
The first path is upstream: a webhook sent twice, an event republished after a lost acknowledgement, a poller rereading an overlapping window, or a producer that does not retain its identifier. The two runs then have separate trigger occurrences but the same business key.
The second path is debatching. With splitOn, every array item starts a separate workflow run. Inspect the exact payload and split expression. If two array elements carry the same key, the platform deliberately creates two runs; reducing concurrency does not merge them.
The third path is overlapping execution. Two valid occurrences can run concurrently, read the same unchanged state, and both write. Trigger concurrency set to one may serialize this workflow, but it does not protect against later redelivery or a second producer instance.
The fourth path is an ambiguous retry. An HTTP action times out from the workflow’s point of view after the target has committed the write. A retry then sends the same intent again. Disabling retries removes that second attempt, but it also turns transient failures and 429 responses into loss or manual recovery. The durable control belongs at the boundary that creates the effect.
Read retries as attempts, not runs
Open both runs and find the first divergence. If each run contains one successful attempt, investigate the trigger and producer. If one run contains multiple attempts for an action after a timeout, 408, 429, or 5xx response, correlate every attempt with target-side logs. A failed orchestrator status does not prove that the server produced no effect.
Capture the effective retry policy, including connector defaults. Compare the client timeout with target processing time. An API that commits and answers late creates an ambiguity window; only increasing the timeout moves that window without making the operation idempotent.
Enforce idempotency at the business boundary
The target should recognize an intent it has already processed and return the original result without creating the effect again. The key must remain stable for the same business operation and change for a new deliberate operation. A useful shape is orderId + operation + version, not a timestamp or run ID.
{
"eventId": "order-84721",
"operation": "create-fulfillment",
"entityVersion": 3,
"idempotencyKey": "order-84721:create-fulfillment:v3",
"correlation": {
"workflowRun": "run-a",
"action": "Create_fulfillment"
}
} The barrier must be atomic. A separate “does this key exist?” read followed by a write is still vulnerable to concurrent runs. Prefer conditional creation, a unique constraint, or a transaction that stores the key and effect together. Retain the result, status, and a lifetime that covers the maximum redelivery window.
If the target cannot change immediately, place a shared idempotency barrier before the action. Treat it as a compensating control: storage failure, premature expiry, a stuck in_progress state, and crash recovery all need defined behavior.
absent -> in_progress -> completed
-> failed_retryable
-> failed_final
On duplicate key
completed return stored result; do not write again
in_progress wait or return accepted; do not start a second write
failed_retryable retry with the same key and bounded ownership
failed_final stop and require an explicit decision Canary without creating a third effect
First deploy key propagation and observation without removing retries. For a controlled test order, deliver the same event twice and, when the test environment permits it, delay the response after a successful write. The expected result is one downstream effect, two correlated traces, and a stable response to the second request.
Then run the negative test: two valid operations on the same entity with different versions or intents must both proceed. An overly broad key eliminates duplicates by suppressing legitimate changes as well.
During the canary, count events, trigger occurrences, runs, attempts, and downstream effects by key. These counts do not have to match: splitOn and retries can increase intermediate levels. The invariant is one business effect per idempotent intent.
Decide, validate, or roll back
Fix the source when it emits duplicate events without a valid reason. Correct splitOn when the array or expression produces unexpected work units. Reduce concurrency when the target is saturated, but do not present serialization as a uniqueness guarantee. Keep retries for transient failures when the downstream boundary can safely absorb repetition.
Roll back if the new key merges distinct operations, if the idempotency store becomes a failure point, or if replayed responses do not match the original result. Restore the previous workflow revision, block resubmissions, and keep a bounded manual procedure while correcting the design. Do not delete idempotency records already created; they still protect against in-flight events.
Final validation needs three proofs: replaying the same event creates no new effect, a new legitimate intent is still accepted, and a timeout after commit converges on the existing result. Record the key retention period, its owner, and the recovery procedure for a stuck state.
Conclusion
A duplicate Logic Apps run is not automatically a broken retry. It can originate in the producer, splitOn, overlapping runs, or a response lost after a real effect. The identifier chain reveals the first level that duplicates before several settings are changed at once.
The runbook ends with an observable decision: repair the producer, correct debatching, bound concurrency, make the target idempotent, or roll back the barrier. Resilience remains valuable when repeating an attempt no longer means repeating the business effect.