AI

AgentOps: validate regional failover before rerouting a production agent

A production runbook for qualifying an AI agent regional failover across model deployment, state, retrieval, tools, identity, traces, canary traffic, validation and rollback.

04 Sept 2026 azureaiagentopsagentsmicrosoft-foundryresiliencefailoverobservabilityevaluationautomationguardrailsrunbookrollbackproduction

An operations agent starts timing out during an incident. The primary model deployment returns intermittent errors, latency rises, and the traffic layer can send requests to a second Azure region. Rerouting looks like the obvious recovery action. It can also move the agent to a different model revision, an older retrieval index, another identity boundary or a tool endpoint that has never executed the production runbook.

This is not a model endpoint failover alone. The production unit is the whole agent path: instructions, model deployment, conversation state, retrieval, tool contracts, identities, network access, approval policy and telemetry. This runbook decides whether to hold, canary the secondary path, fail over a bounded workload or return traffic to the primary region.

Freeze the failing path and the recovery objective

Start with one workload and one time window. Do not mix interactive diagnosis, background evaluation and write-capable automation in the same failover decision.

yaml regional-failover-incident.yml
incident:
id: INC-2841
detected_at: 2026-09-04T07:42:00Z
agent: operations-assistant
request_class: interactive-diagnosis
primary_region: region-a
candidate_region: region-b
symptom: timeout-and-upstream-errors

recovery_objective:
restore: read-only incident diagnosis
preserve:
- approved knowledge boundary
- refusal and approval policy
- trace correlation
- tool-call idempotency
excluded_initially:
- production writes
- long-running conversations already in execution
- batch evaluations

decision_deadline: 20m
owner: platform-oncall

The objective is deliberately narrower than “restore the agent.” Read-only diagnosis can often move first. A production action should wait until state, tool and approval guarantees have been proved on the candidate path.

Define the real failover unit

A healthy secondary model endpoint does not prove that the agent is healthy. Inventory every dependency crossed by one representative request and mark whether it is regional, shared or pinned.

text agent-failover-unit.txt
Layer                     Primary path       Candidate path      Required proof
Agent instructions         version 42         version 42           immutable digest
Model deployment           deployment-a       deployment-b         approved model and config
Conversation state         state-a            replicated/shared    consistency and ownership
Retrieval index            index-a-118         index-b-118          corpus and filter parity
Tool registry              tools-a-17          tools-b-17           schema and policy digest
Runtime identity           identity-a          identity-b           least-privilege access
Network and DNS            route-a             route-b              read and tool reachability
Approval store             shared              shared               atomic single use
Telemetry                  workspace/app        workspace/app        trace continuity
Kill switch                route policy        route policy         tested return path

“Same configuration” is not enough. Record deployable identifiers and digests. If the secondary region cannot prove which instructions, tool schema or index snapshot it serves, keep it outside the write path.

Qualify the secondary region without production side effects

Run synthetic checks from the same network and identity class used by the agent. Validate the model call, retrieval filters, read-only tools, trace export and policy refusal independently before testing an end-to-end conversation.

yaml secondary-region-probes.yml
probes:
- id: model-minimal-response
  input: fixed-health-prompt
  assert: response-and-trace

- id: retrieval-approved-source
  input: known-document-id
  assert: expected-version-and-access-filter

- id: tool-read-only
  tool: get_service_health
  assert: schema-valid-and-no-write

- id: forbidden-production-write
  tool: restart_service
  approval: absent
  assert: rejected-before-dispatch

- id: trace-chain
  assert:
  - region-and-deployment-recorded
  - model-and-tool-spans-correlated
  - sensitive-content-policy-applied

stop_on:
- unknown-configuration-version
- missing-refusal
- missing-trace
- identity-scope-wider-than-primary

Do not use a destructive tool as a health probe. The negative test should prove that the policy gate refuses a write before dispatch, not that the backend can undo it afterward.

Replay outcomes, not wording

Before sending live traffic, replay a fixed evaluation set against both regions. Compare task outcome, grounding, tool selection, arguments, refusal behavior and latency envelope. Exact wording is a weak parity signal.

json regional-parity-case.json
{
"case_id": "diagnose-api-latency-07",
"input_snapshot": "eval-set-20260904-v3",
"expected": {
  "allowed_sources": ["runbook-api-v8", "service-map-v12"],
  "required_tool_sequence": ["search_runbook", "query_health_readonly"],
  "forbidden_tools": ["restart_service", "change_route"],
  "required_decision": "collect-more-evidence",
  "required_trace_fields": ["region", "deployment", "agent_version", "tool_call_id"]
},
"compare": [
  "task_completion",
  "groundedness",
  "tool_call_accuracy",
  "refusal_consistency",
  "latency",
  "token_usage"
]
}

Use representative operational cases, including ambiguous requests, denied writes, unavailable tools and stale knowledge. A secondary path that answers fluently but chooses a broader tool action has failed the readiness test.

Protect state and side effects during the switch

Failover becomes dangerous when a timed-out request may still be running in the primary region. Retrying the whole turn in the secondary region can duplicate a tool call or consume the same approval twice.

text failover-state-rules.txt
Before retrying a conversation
Read the last committed turn and active run state
Query backend operations by idempotency key
Preserve conversation, trace and tool-call correlation IDs
Reject concurrent ownership of the same run

Safe to replay
Model call with no dispatched tool
Retrieval query against an immutable snapshot
Read-only tool with bounded retry semantics

Do not replay automatically
Write tool with an unknown backend result
Consumed or pending human approval
Long-running action still owned by the primary worker
Turn whose state version cannot be reconciled

Recovery for ambiguous writes
Query operation state
Reconcile observed target state
Resume by idempotency key or compensate
Never infer failure from a client timeout alone

The traffic router should not decide these rules. It can select a healthy endpoint, but the orchestration layer must own run state, idempotency and approval consumption.

Canary a bounded traffic class

Move a small, identifiable class of new read-only sessions first. Keep session affinity during each run so one conversation does not alternate regions while tools or retrieval state differ.

yaml agent-failover-canary.yml
canary:
eligible:
- new-session
- read-only-diagnosis
- approved-evaluation-cohort
excluded:
- active-production-action
- pending-human-approval
- unreconciled-primary-run

candidate_weight_percent: 5
session_affinity: required
observation_window: 15m

promotion_gates:
operational:
- run-success-rate-within-approved-envelope
- latency-within-approved-envelope
- no-unbounded-retry-growth
quality:
- task-and-grounding-gates-pass
- tool-selection-gate-passes
control:
- refusal-parity-passes
- trace-completeness-passes
- no-duplicate-tool-operation

automatic_stop:
- write-dispatched-without-valid-approval
- state-ownership-conflict
- missing-region-or-version-in-trace
- retrieval-snapshot-mismatch

Set thresholds from the service’s observed baseline and risk policy rather than copying generic percentages. One control breach should stop the canary even when average latency improves.

Correlate the failover in traces

Microsoft Foundry tracing can feed agent execution telemetry into Application Insights. Whatever instrumentation path is used, the failover investigation needs a continuous chain across request, model, retrieval and tools.

kusto 01-agent-regional-failover.kql
let IncidentStart = datetime(2026-09-04T07:42:00Z);
let IncidentEnd = IncidentStart + 45m;
AgentRunEvents
| where TimeGenerated between (IncidentStart .. IncidentEnd)
| where AgentName == "operations-assistant"
| summarize
  Runs = dcount(RunId),
  FailedRuns = dcountif(RunId, RunStatus != "completed"),
  P95LatencyMs = percentile(DurationMs, 95),
  ToolCalls = countif(EventType == "tool_call"),
  DuplicateOperations = dcountif(OperationId, IsDuplicate == true),
  MissingTraceContext = countif(isempty(TraceId) or isempty(DeploymentVersion))
by Region, DeploymentName, AgentVersion, bin(TimeGenerated, 5m)
| extend FailureRate = todouble(FailedRuns) / Runs
| order by TimeGenerated asc

Adapt the table and field names to the telemetry contract. Pair operational metrics with sampled evaluations: low error rates do not reveal grounding drift, changed tool choices or weakened refusals.

Decide hold, fail over, fail back or rollback

The runbook ends with an explicit traffic and control decision.

text regional-failover-decision.txt
Hold on the primary path
Primary degradation is not reproduced or dependency scope is unknown
Candidate configuration, state or traces cannot be proved
Recovery risk is greater than the bounded incident impact

Fail over read-only traffic
Candidate probes and evaluation gates pass
New sessions can be isolated and traced
Write tools remain blocked

Promote the candidate path
Canary meets operational, quality and control gates
State ownership and idempotency are proven
Tool, identity and approval boundaries match the approved contract

Fail back
Primary health is stable through the observation window
New sessions return gradually with session affinity
Existing candidate runs finish or are explicitly reconciled

Rollback the routing change
Quality, refusal, trace or state gate fails
Candidate weights return to zero
Ambiguous actions are reconciled before any retry
Incident evidence retains both regional paths

Automate probe collection, traffic weights and stop conditions where their semantics are deterministic. Keep the production-write decision behind an explicit gate until the secondary path has proved the same action boundary as the primary.

Conclusion

Regional failover for an AI agent is a change to an operational system, not a DNS shortcut around an unhealthy model. The model may recover while state, retrieval, tools, identities or approval controls diverge.

Fail over the smallest useful workload, prove versions and policies, evaluate outcomes, protect side effects, and canary with trace continuity. Promote only when availability, quality and control pass together. Otherwise hold or return traffic, reconcile every ambiguous action, and keep the rollback narrower than the incident.