Cloud

Azure Service Bus: diagnose message-lock loss before extending auto-renewal

A production runbook for separating processing time, connection loss, renewal failure and late settlement before extending an Azure Service Bus message lock.

28 Sept 2026 azureservice-busmessagingpeeklockmessage-lockidempotencyobservabilityautomationincidentcanaryrunbookrollbackproduction

An Azure Service Bus worker finishes a payment reconciliation, then fails while completing the message with a lock-lost exception. The message returns to the queue, another instance receives it, and the business action may run twice. Extending automatic lock renewal looks like the quickest fix. It can also hide a slow dependency, a blocked event loop, connection churn or a handler whose side effects are not idempotent.

The running case is a PeekLock consumer on payments-reconcile-prod. Most messages finish in 20 seconds, but a small cohort takes several minutes. The runbook separates the initial entity lock, client renewal, processing and settlement. It ends with a bounded decision: shorten or split the work, adjust renewal, repair connectivity, quarantine a message family, or roll back the worker.

Freeze one delivery before changing the lock

Capture one failed delivery from receive to settlement. Record the namespace, entity, message ID, delivery count, sequence number, worker instance, deployment version, receive time, LockedUntil, business-operation ID, completion time and exact exception.

Do not replay the message or increase LockDuration yet. The first question is whether the handler lost the lock before, during or after the business side effect.

yaml message-lock-incident.yml
entity: payments-reconcile-prod
receive_mode: PeekLock
message_id: pay-20260928-1842
delivery_count: 3
worker_instance: reconcile-7d8f9
release: 2026.09.28.4
received_utc: 2026-09-28T14:02:11Z
locked_until_utc: 2026-09-28T14:03:11Z
business_commit_utc: 2026-09-28T14:04:37Z
settlement_utc: 2026-09-28T14:04:38Z
result: MessageLockLost
side_effect_key: reconciliation/pay-20260928-1842

A lock-lost exception does not prove that the business action failed. It proves that the receiver could not settle with the lock token it held. Treat the external side effect and the broker settlement as two separate facts.

Read the entity contract

Inspect the deployed queue or subscription instead of relying on an application setting or portal memory.

bash 01-read-service-bus-lock-contract.sh
RG="rg-messaging-prod"
NAMESPACE="sb-orders-prod"
QUEUE="payments-reconcile-prod"

az servicebus queue show --resource-group "$RG" --namespace-name "$NAMESPACE" --name "$QUEUE" --query "{lockDuration:lockDuration,maxDeliveryCount:maxDeliveryCount,deadLetterOnExpiration:deadLetteringOnMessageExpiration,status:status}" --output yaml

In PeekLock, the entity grants an initial lock. Completion, abandonment, deferral and dead-lettering settle the message while that lock is valid. If the lock expires, the message becomes available for redelivery. Repeated delivery can eventually move it to the dead-letter queue according to MaxDeliveryCount.

A longer entity lock is not free. When a worker dies, another worker must wait longer before the message becomes available again. Keep the initial lock longer than normal processing, but use renewal for genuinely long, bounded work.

Separate four failure families

A useful diagnosis distinguishes:

  1. processing overrun: handler duration exceeds the lock or the configured auto-renewal window;
  2. renewal starvation: thread-pool exhaustion, event-loop blocking, CPU pressure or a frozen process prevents renewals from running on time;
  3. connection or receiver loss: the AMQP link, receiver or client is closed, recreated or disconnected before settlement;
  4. late settlement: the work completed, but the handler settles after cancellation, shutdown or lock expiry.

Entity property changes, service updates and connection loss can also invalidate volatile locks. Increasing only the renewal duration is justified only after proving that renewal remains healthy and processing is intentionally longer.

Measure the processing envelope

Instrument receive, first side effect, last side effect and settlement as separate spans or structured events. Include MessageId, DeliveryCount, LockedUntil, worker instance and release.

kusto 02-message-lock-processing-envelope.kql
let Window = 6h;
AppTraces
| where TimeGenerated > ago(Window)
| where Properties["Entity"] == "payments-reconcile-prod"
| extend
  MessageId = tostring(Properties["MessageId"]),
  Phase = tostring(Properties["Phase"]),
  DeliveryCount = toint(Properties["DeliveryCount"]),
  Worker = tostring(Properties["WorkerInstance"]),
  Release = tostring(Properties["Release"])
| summarize
  Received=minif(TimeGenerated, Phase == "received"),
  BusinessCommitted=maxif(TimeGenerated, Phase == "business_committed"),
  Settled=maxif(TimeGenerated, Phase == "settled"),
  LockLost=countif(Phase == "lock_lost"),
  MaxDelivery=max(DeliveryCount)
by MessageId, Worker, Release
| extend
  ProcessingSeconds=datetime_diff("second", BusinessCommitted, Received),
  SettlementLagSeconds=datetime_diff("second", Settled, BusinessCommitted)
| order by LockLost desc, ProcessingSeconds desc

Compare p50, p95 and p99 processing time with the initial lock and the maximum renewal duration. Then split by message type, dependency, release and worker instance. A long tail isolated to one document type calls for workload shaping; a whole instance with missing renewals points to runtime health.

Platform metrics provide another boundary. Compare completed and abandoned messages, active and dead-lettered backlog, server errors and connection churn for the same UTC window. Metrics show the entity trend; application traces explain which handler owned each lock.

Prove whether renewal actually runs

Read the effective SDK configuration at startup and expose it in deployment diagnostics. For the current .NET client, MaxAutoLockRenewalDuration bounds how long the processor renews locks automatically. It is not the duration of each lock extension and it does not make the handler idempotent.

csharp service-bus-processor-options.cs
var options = new ServiceBusProcessorOptions
{
  ReceiveMode = ServiceBusReceiveMode.PeekLock,
  AutoCompleteMessages = false,
  MaxConcurrentCalls = 8,
  MaxAutoLockRenewalDuration = TimeSpan.FromMinutes(10)
};

logger.LogInformation(
  "ServiceBus processor configured: mode={Mode}, concurrent={Concurrent}, renewal={Renewal}",
  options.ReceiveMode,
  options.MaxConcurrentCalls,
  options.MaxAutoLockRenewalDuration);

Do not copy ten minutes as a universal value. Set the renewal window from a measured processing budget, with margin, and alert when work approaches it. If a task can run without a credible upper bound, move the long operation behind a durable workflow and let the Service Bus handler persist intent quickly.

Also inspect cancellation and disposal paths. A deployment that stops the processor before in-flight settlement can produce a lock-loss cluster even when steady-state processing is fast.

Make redelivery safe before tuning

PeekLock provides at-least-once handling, not exactly-once business effects. Use a stable business key or message ID to guard the external action. Store an operation state such as started, committed and settled-pending in a system that supports the required consistency.

On redelivery:

  • if no operation exists, claim the key and process;
  • if the operation is committed, skip the side effect and complete the message;
  • if the state is ambiguous, quarantine or reconcile instead of repeating the action;
  • if the prior attempt failed before commitment, resume only from a known boundary.

Do not switch to ReceiveAndDelete to remove lock exceptions. That exchanges duplicates for possible message loss if the worker fails after receive.

Canary one correction

Choose one message family or a dedicated canary queue. Keep production concurrency unchanged and modify one variable: handler decomposition, renewal duration, dependency timeout or worker capacity.

The canary must prove:

  1. a normal message completes once and settles before the initial lock expires;
  2. a controlled long message renews successfully and completes once;
  3. a killed worker causes safe redelivery without duplicating the business effect;
  4. a dependency timeout ends before the renewal budget is exhausted;
  5. deployment shutdown drains or abandons in-flight work predictably.

Observe several lock periods and at least one controlled failure. Compare processing percentiles, renewal errors, completion, abandonment, delivery count and dead-letter growth with the previous release.

Decide, validate or roll back

Extend auto-renewal only when long processing is expected, bounded, observable and idempotent, and when renewal continues to run under realistic CPU and connection conditions. Shorten or split the handler when it holds a broker lock while waiting for work that can be persisted elsewhere. Repair runtime or network health when renewals disappear across unrelated messages on the same instance.

Roll back when the canary increases duplicate side effects, delivery count, dead-letter growth or shutdown time. Restore the previous processor settings and release, stop the candidate consumers, then verify that one known message is received, processed and settled once. Keep ambiguous deliveries quarantined until their business state is reconciled.

Conclusion

A lost Service Bus message lock is a timeline problem before it is a duration problem. Reconstruct receive, renewal, side effect and settlement; compare the processing tail with the deployed lock contract; and prove that redelivery is safe.

The operable outcome is explicit: split slow work, repair renewal, extend a measured window, quarantine an unsafe cohort or return to the last known worker. The incident is closed only when broker state and business state agree.