AI

AgentOps: stop a multi-agent handoff loop before actions multiply

A production runbook for detecting an agent delegation loop, containing side effects, rebuilding traces, and deciding controlled recovery or a single-agent rollback.

12 Aug 2026 aiagentopsagentsmulti-agenthandofforchestrationobservabilitytracesguardrailsidempotencyautomationrunbookrollbackproduction

A triage agent hands a network incident to a cloud specialist. The cloud agent asks an automation agent for a check, which sends the task back to triage because approval is missing. The conversation appears active, yet the same work starts again under new run identifiers. Tool calls grow, several tickets are prepared, and a retry may eventually execute the same action twice.

This is more than an answer-quality issue. A handoff loop is an orchestration incident: task state, ownership and previous side effects no longer travel reliably with delegation. This runbook follows an internal operations agent system that reads runbooks, queries logs and prepares actions for approval. The goal is to stop multiplication without destroying evidence, then decide whether multi-agent routing can recover or must temporarily return to a single owner.

Recognize a loop before limits are exhausted

A legitimate delegation changes ownership because a capability is missing. A loop repeats the same intent without reducing uncertainty, adding evidence or moving the task toward a terminal state.

text handoff-loop-signals.txt
Signals to correlate
same task_id or business fingerprint across multiple traces
return to a previously visited agent without new evidence
handoff depth or count exceeds the budget
same tool called with equivalent arguments
new tickets or jobs prepared for the same intent
retries while the previous result remains unknown
latency and cost rise without a state transition

A user-facing timeout is a late signal. The platform should observe depth, visited agents, action fingerprints and the last useful result. Different prose is not progress when it prepares the same operation.

Freeze side effects, preserve evidence

Once a loop is likely, block new handoffs and write-capable tools for the affected scope. Keep read-only inspection, traces and diagnostics available. Do not purge queues or conversations before identifying every prepared, approved or executed action.

yaml handoff-containment.yml
incident_scope:
workflow: incident-triage
task_id: task-7f31
correlation_id: corr-91ab

containment:
new_handoffs: blocked
write_tools: prepare_only
automatic_retries: disabled
read_only_diagnostics: allowed
trace_retention: preserved

operator_checks:
- list prepared and approved actions
- cancel duplicate pending work
- identify any executed side effect
- assign one human incident owner

Containment should target a task, workflow or action class rather than every agent when possible. If the orchestrator cannot isolate that boundary, force sensitive tools into read-only mode globally until in-flight work is qualified.

Rebuild the chain of responsibility

An operable trace links the logical task to runs, handoffs and tools. Each transfer records who delegates, who receives ownership, why, from which state and for which expected result. Without that envelope, two runtimes can recreate the same task and bypass a local counter.

json handoff-envelope.json
{
"taskId": "task-7f31",
"correlationId": "corr-91ab",
"parentRunId": "run-triage-04",
"runId": "run-cloud-05",
"fromAgent": "incident-triage",
"toAgent": "cloud-diagnostics",
"reasonCode": "missing-network-evidence",
"expectedOutcome": "read-only route evidence",
"visitedAgents": ["incident-triage", "cloud-diagnostics"],
"hop": 2,
"hopBudget": 4,
"actionFingerprint": "diagnose:subscription:resource:route",
"stateVersion": 6,
"deadline": "<timestamp>",
"sideEffectMode": "none"
}

taskId remains stable for the business intent; runId changes for each execution. actionFingerprint deduplicates a logical operation even when prompt wording or argument order changes. stateVersion lets the orchestrator reject a write based on stale context.

Read the loop from traces

The physical schema depends on the platform, but the query should group events by logical task and rebuild the agent sequence. Adapt the example table and columns to the telemetry actually emitted.

kusto 01-detect-handoff-loops.kql
let Window = 2h;
AgentOrchestrationEvents
| where TimeGenerated > ago(Window)
| where EventType in ("handoff", "tool_call", "terminal_state")
| summarize
  FirstSeen = min(TimeGenerated),
  LastSeen = max(TimeGenerated),
  Handoffs = countif(EventType == "handoff"),
  Agents = make_list_if(ToAgent, EventType == "handoff", 32),
  DistinctAgents = dcountif(ToAgent, EventType == "handoff"),
  RepeatedActions = count() - dcount(ActionFingerprint),
  TerminalStates = countif(EventType == "terminal_state")
by TaskId, CorrelationId
| extend SuspectedLoop = Handoffs > 4
  or (Handoffs > DistinctAgents and TerminalStates == 0)
  or RepeatedActions > 1
| where SuspectedLoop
| order by Handoffs desc, LastSeen desc

Do not conclude from handoff count alone. A complex workflow may legitimately cross several specialists. Strong evidence combines a revisit, no new state, a repeated action and no terminal outcome.

Enforce the orchestration contract outside the prompt

Telling agents not to loop is insufficient. The orchestrator must reject an invalid transition before invoking another runtime.

yaml orchestration-guardrails.yml
handoff_policy:
max_hops: 4
revisit_agent: require_new_evidence
deadline: required
terminal_states: [resolved, refused, needs_human, failed]

progress_policy:
require_one_of:
  - new_evidence_id
  - narrower_scope
  - validated_state_transition
reject_same_action_fingerprint: true

side_effect_policy:
idempotency_key: required
stale_state_write: reject
write_after_handoff: require_fresh_approval
unknown_tool_result: never_retry_automatically

The hop budget bounds execution, but idempotency protects production. Every write tool should accept a key derived from the task and logical action, then return the known result instead of starting again. If a call times out with an unknown outcome, query execution status before any retry.

Fix the cause, not only the counter

A loop usually exposes an incomplete contract: overlapping ownership, free-form handoff reasons, unshared state, no needs_human outcome, or a tool that does not separate preparation from execution. Classify the incident before tuning prompts.

When two agents claim the same domain, define one ownership rule. When neither can finish because evidence is missing, add an explicit human terminal state. When context is lost, centralize logical task state and require conditional writes against stateVersion. When a tool runs twice, repair idempotency before reopening handoffs.

Replay without real actions

Reproduce the failing trace with simulated or read-only tools. Tests should verify progression, terminal states and duplicate prevention rather than the final answer alone.

yaml handoff-regression-cases.yml
cases:
- name: missing_evidence_returns_to_previous_agent
  expected:
    - second_visit_requires_new_evidence
    - otherwise_needs_human
    - no_write_tool

- name: tool_timeout_with_unknown_result
  expected:
    - query_execution_status
    - no_automatic_retry
    - same_idempotency_key

- name: hop_budget_exhausted
  expected:
    - terminal_state_needs_human
    - full_handoff_summary
    - pending_duplicates_cancelled

promotion_gate:
repeated_action_fingerprints: 0
writes_from_stale_state: 0
unbounded_tasks: 0
trace_completeness: required

Include a valid case where several agents cooperate and reach completion. A guardrail that prevents every delegation removes the loop and the value of the architecture with it.

Decide recovery or a single-agent rollback

Restore read-only diagnostics first on a small scope. Then open handoffs with an enforced budget and complete traces, followed by preparation tools. Writes return only after idempotency, approval and timeout cases pass.

Roll back temporarily to one agent when shared state remains ambiguous, tools cannot guarantee idempotency, or an executed action cannot be linked to a task. The rollback assigns one owner, disables the affected multi-agent routes and retains the same tool controls. It must not turn the single agent into a shortcut around bounded writes.

Conclusion

A handoff loop should be handled like a distributed-systems failure: stable task identity, versioned state, deadlines, hop budgets, terminal outcomes and idempotent effects. Traces must prove progress, not merely count tokens.

The final decision is operational: recover progressively when every transfer adds evidence and duplicate actions are impossible, or restore a single-agent path until orchestration guarantees those properties. The goal is not more delegation. It is delegation that remains explainable and reversible.