Cloud
Azure Service Bus: diagnose the dead-letter queue before a bounded replay
A production runbook for classifying dead-lettered messages, proving the cause, checking idempotency and replaying with a canary without creating a second incident.
The dead-letter queue for an Azure Service Bus queue or subscription is growing. The consumer is still running, but some messages exceed their maximum delivery count, expire or are rejected by application logic. Sending everything back to the main entity looks like the quickest way to restore the flow. In production, that bulk replay can reintroduce an invalid payload, overload the dependency that is already failing and duplicate business effects that partially succeeded.
The running case is a subscription that distributes orders to a processing service. A recent release changed a contract, while a downstream API also experienced transient errors. This runbook separates those two populations, fixes the cause and allows only a replay with explicit scope, rate and stop conditions.
Pin the entity and incident window
A DLQ belongs to one specific queue or subscription. It is not a namespace-wide queue. Start with the full entity path, consumer, first visible increase and changes immediately before it. For a topic, inspect each subscription: the same published message can be healthy for one subscription and fail in another.
incident: inc-20260816-011
namespace: <service-bus-namespace>
entity_type: subscription
topic: <topic-name>
subscription: <subscription-name>
environment: production
first_dlq_growth_utc: <timestamp>
last_known_healthy_utc: <timestamp>
consumer:
application: <service-name>
version: <release-id>
identity: <managed-identity-or-sas-name>
receive_mode: peek-lock
preserve:
- active and dead-letter message counts
- incoming and outgoing message rates
- throttled requests, server errors and user errors
- consumer exceptions, lock losses and dependency failures
- deployments, configuration and subscription rule changes
- a redacted sample grouped by dead-letter reason Freeze deployments and manual purge operations. Do not record only the current count: preserve a time series. A stable DLQ containing old messages does not require the same response as a slope that is still increasing.
Measure without consuming
Start with the entity counters, then open the dead-letter subqueue for inspection. A peek operation neither locks nor removes messages. It produces a sample, not perfect statistical evidence, so inspect multiple positions when the population is heterogeneous.
az servicebus topic subscription show \
--resource-group <resource-group> \
--namespace-name <namespace> \
--topic-name <topic> \
--name <subscription> \
--query '{active:countDetails.activeMessageCount,deadLetter:countDetails.deadLetterMessageCount,transferDeadLetter:countDetails.transferDeadLetterMessageCount}' For each sampled message, collect these fields without exposing the full payload: MessageId, business or correlation identifier, SequenceNumber, enqueue time, delivery count, type/schema, DeadLetterReason, DeadLetterErrorDescription and relevant application properties. Hash or mask sensitive values.
receiver = client.get_subscription_receiver(
topic_name=topic,
subscription_name=subscription,
sub_queue=ServiceBusSubQueue.DEAD_LETTER,
)
sample = await receiver.peek_messages(max_message_count=50)
for message in sample:
emit_redacted({
"message_id": message.message_id,
"sequence_number": message.sequence_number,
"delivery_count": message.delivery_count,
"enqueued_time_utc": message.enqueued_time_utc,
"dead_letter_reason": message.dead_letter_reason,
"dead_letter_description": message.dead_letter_error_description,
"schema": message.application_properties.get(b"schema"),
}) Inspect the transfer DLQ as well when the architecture uses autoforwarding. A transfer failure remains attached to the source entity and should not be diagnosed as a final consumer failure.
Classify before fixing
Group messages by reason, schema version, type, time window and release identifier. Avoid collapsing the incident into “the consumer is broken.”
MaxDeliveryCountExceeded
Look for consumer exceptions, expired locks, repeated abandon calls or a slow downstream dependency.
Do not raise MaxDeliveryCount until another attempt has a proven path to success.
TTLExpiredException
Compare TTL, active wait time, backlog and consumer capacity.
An expired message may be functionally obsolete; replay is not automatic.
Application dead-letter
Read the reason and description emitted by the handler.
Separate contract errors, business rules, authentication and transient failure.
Filter evaluation or session
Review subscription rules, the required SessionId and the contract change.
Fix routing before returning the message to circulation.
MaxTransferHopCountExceeded or transfer DLQ
Rebuild the autoforward chain and destination state.
Do not replay into a destination whose path remains invalid. This classification should produce homogeneous batches. One batch may be replayable after a dependency recovers; another needs payload migration; a third should remain quarantined because the command is no longer valid.
Prove the cause outside the DLQ
The dead-letter reason describes the final mechanism, not necessarily the initial cause. MaxDeliveryCountExceeded can result from an invalid payload, a failing downstream API, a lock that is too short or a consumer that performs the work and fails before complete.
Build a timeline from four sources: Service Bus metrics, consumer logs, dependency traces and platform changes. Look for:
- downstream errors rising with abandon operations;
MessageLockLosterrors after processing exceeds the lock window;- one release or schema version dominating the DLQ;
- throttled requests at namespace level;
- business effects written before settlement failed;
- a changed subscription rule or autoforward chain.
The deadLetterMessageCount is a stock measurement. Request metrics and application logs explain the flow feeding it. Do not decide from one chart.
Verify idempotency and downstream state
Before replaying anything, take a few MessageId values and find them in the database, downstream API, traces and audit log. A Service Bus consumer using peek-lock has at-least-once delivery semantics: processing may succeed and the subsequent complete may fail. The message then returns even though its effect already exists.
Document for each class:
- the actual idempotency key, preferably a stable business identifier;
- the observable effect and how it is compared with the expected result;
- the duplicate policy: ignore, update, compensate or block;
- compatibility between the old payload and the corrected handler;
- the business date after which the command must no longer execute.
Service Bus duplicate detection helps with some producer retries. It does not replace consumer idempotency, and it can discard a replay that reuses a MessageId still covered by the detection window.
Prepare a replay manifest
Do not turn the DLQ into a “send all” button. Produce a versioned manifest that names the eligible messages and exclusions precisely.
incident: inc-20260816-011
source: <topic>/<subscription>/$DeadLetterQueue
selection:
dead_letter_reason: MaxDeliveryCountExceeded
enqueued_from_utc: <timestamp>
enqueued_to_utc: <timestamp>
schema_versions: [order.v3]
exclude_business_states: [cancelled, already_completed]
controls:
dry_run: true
canary_messages: 1
batch_size: 20
max_messages: 200
max_rate_per_second: 2
pause_between_batches_seconds: 60
approval_required: true
stop_when:
- any duplicate side effect
- consumer error rate above agreed threshold
- downstream latency above agreed threshold
- dead-letter count grows from replayed messages
- identity, schema or target cannot be proven Archive the redacted envelope, payload hash, source sequence, decision and result for every message. A replay tool needs a ledger so rerunning the runbook cannot silently resend the same batch.
Run a canary, then bounded batches
The first pass should be a dry run: the handler validates schema, identity, dependencies and expected effect without a business write. Then replay one representative message. Wait for consumption, the downstream effect and the absence of a new dead-letter before authorizing the first batch.
When moving a message, complete the source only after confirming the send and recording it in the ledger. If the send succeeds but completing the DLQ message fails, the source can remain visible. The ledger and consumer idempotency must make that ambiguity harmless.
Between batches, inspect the source DLQ, active messages, processing rate, errors, throttling, downstream latency and business duplicates. Replay rate must stay below the measured spare capacity of the consumer and dependency, not merely the theoretical capacity of Service Bus.
Decide, validate or roll back
The runbook ends with an explicit decision:
REPLAY
Cause fixed and proven.
Homogeneous batch, compatible payload and idempotent effect.
Canary passed, stop conditions active and an owner present.
TRANSFORM THEN REPLAY
Old schema understood and transformation tested outside production.
Original payload archived, mapping reviewable and canary mandatory.
KEEP IN QUARANTINE
Downstream state ambiguous, message obsolete or business rule unresolved.
Preserve evidence and decision; do not purge to lower a counter.
DISCARD WITH APPROVAL
Message explicitly non-executable and retention policy validated.
Record identifiers, reason, approver and expected impact.
ROLL BACK REPLAY
Stop the replay worker and new batches.
Leave remaining messages in the DLQ.
Block the failing class, verify canary effects and compensate if required. Final validation is not a zero DLQ count. It is evidence that the growth rate returned to normal, eligible messages produced exactly the expected effect, exclusions remain justified and a new deployment does not recreate the same failure class.
Conclusion
A dead-letter queue is an operational boundary, not a waste bin. It holds messages the system could not process safely. The right response is to classify, correlate, verify downstream state and bound the replay.
With a manifest, canary, ledger and stop conditions, the team can return only proven-safe messages to circulation. Rollback stays simple: stop replay, keep the remainder quarantined and fix the cause before resuming.