AI
AgentOps: validate regional failover before rerouting a production agent
A production runbook for qualifying an AI agent regional failover across model deployment, state, retrieval, tools, identity, traces, canary traffic, validation and rollback.
An operations agent starts timing out during an incident. The primary model deployment returns intermittent errors, latency rises, and the traffic layer can send requests to a second Azure region. Rerouting looks like the obvious recovery action. It can also move the agent to a different model revision, an older retrieval index, another identity boundary or a tool endpoint that has never executed the production runbook.
This is not a model endpoint failover alone. The production unit is the whole agent path: instructions, model deployment, conversation state, retrieval, tool contracts, identities, network access, approval policy and telemetry. This runbook decides whether to hold, canary the secondary path, fail over a bounded workload or return traffic to the primary region.
Freeze the failing path and the recovery objective
Start with one workload and one time window. Do not mix interactive diagnosis, background evaluation and write-capable automation in the same failover decision.
incident:
id: INC-2841
detected_at: 2026-09-04T07:42:00Z
agent: operations-assistant
request_class: interactive-diagnosis
primary_region: region-a
candidate_region: region-b
symptom: timeout-and-upstream-errors
recovery_objective:
restore: read-only incident diagnosis
preserve:
- approved knowledge boundary
- refusal and approval policy
- trace correlation
- tool-call idempotency
excluded_initially:
- production writes
- long-running conversations already in execution
- batch evaluations
decision_deadline: 20m
owner: platform-oncall The objective is deliberately narrower than “restore the agent.” Read-only diagnosis can often move first. A production action should wait until state, tool and approval guarantees have been proved on the candidate path.
Define the real failover unit
A healthy secondary model endpoint does not prove that the agent is healthy. Inventory every dependency crossed by one representative request and mark whether it is regional, shared or pinned.
Layer Primary path Candidate path Required proof
Agent instructions version 42 version 42 immutable digest
Model deployment deployment-a deployment-b approved model and config
Conversation state state-a replicated/shared consistency and ownership
Retrieval index index-a-118 index-b-118 corpus and filter parity
Tool registry tools-a-17 tools-b-17 schema and policy digest
Runtime identity identity-a identity-b least-privilege access
Network and DNS route-a route-b read and tool reachability
Approval store shared shared atomic single use
Telemetry workspace/app workspace/app trace continuity
Kill switch route policy route policy tested return path “Same configuration” is not enough. Record deployable identifiers and digests. If the secondary region cannot prove which instructions, tool schema or index snapshot it serves, keep it outside the write path.
Qualify the secondary region without production side effects
Run synthetic checks from the same network and identity class used by the agent. Validate the model call, retrieval filters, read-only tools, trace export and policy refusal independently before testing an end-to-end conversation.
probes:
- id: model-minimal-response
input: fixed-health-prompt
assert: response-and-trace
- id: retrieval-approved-source
input: known-document-id
assert: expected-version-and-access-filter
- id: tool-read-only
tool: get_service_health
assert: schema-valid-and-no-write
- id: forbidden-production-write
tool: restart_service
approval: absent
assert: rejected-before-dispatch
- id: trace-chain
assert:
- region-and-deployment-recorded
- model-and-tool-spans-correlated
- sensitive-content-policy-applied
stop_on:
- unknown-configuration-version
- missing-refusal
- missing-trace
- identity-scope-wider-than-primary Do not use a destructive tool as a health probe. The negative test should prove that the policy gate refuses a write before dispatch, not that the backend can undo it afterward.
Replay outcomes, not wording
Before sending live traffic, replay a fixed evaluation set against both regions. Compare task outcome, grounding, tool selection, arguments, refusal behavior and latency envelope. Exact wording is a weak parity signal.
{
"case_id": "diagnose-api-latency-07",
"input_snapshot": "eval-set-20260904-v3",
"expected": {
"allowed_sources": ["runbook-api-v8", "service-map-v12"],
"required_tool_sequence": ["search_runbook", "query_health_readonly"],
"forbidden_tools": ["restart_service", "change_route"],
"required_decision": "collect-more-evidence",
"required_trace_fields": ["region", "deployment", "agent_version", "tool_call_id"]
},
"compare": [
"task_completion",
"groundedness",
"tool_call_accuracy",
"refusal_consistency",
"latency",
"token_usage"
]
} Use representative operational cases, including ambiguous requests, denied writes, unavailable tools and stale knowledge. A secondary path that answers fluently but chooses a broader tool action has failed the readiness test.
Protect state and side effects during the switch
Failover becomes dangerous when a timed-out request may still be running in the primary region. Retrying the whole turn in the secondary region can duplicate a tool call or consume the same approval twice.
Before retrying a conversation
Read the last committed turn and active run state
Query backend operations by idempotency key
Preserve conversation, trace and tool-call correlation IDs
Reject concurrent ownership of the same run
Safe to replay
Model call with no dispatched tool
Retrieval query against an immutable snapshot
Read-only tool with bounded retry semantics
Do not replay automatically
Write tool with an unknown backend result
Consumed or pending human approval
Long-running action still owned by the primary worker
Turn whose state version cannot be reconciled
Recovery for ambiguous writes
Query operation state
Reconcile observed target state
Resume by idempotency key or compensate
Never infer failure from a client timeout alone The traffic router should not decide these rules. It can select a healthy endpoint, but the orchestration layer must own run state, idempotency and approval consumption.
Canary a bounded traffic class
Move a small, identifiable class of new read-only sessions first. Keep session affinity during each run so one conversation does not alternate regions while tools or retrieval state differ.
canary:
eligible:
- new-session
- read-only-diagnosis
- approved-evaluation-cohort
excluded:
- active-production-action
- pending-human-approval
- unreconciled-primary-run
candidate_weight_percent: 5
session_affinity: required
observation_window: 15m
promotion_gates:
operational:
- run-success-rate-within-approved-envelope
- latency-within-approved-envelope
- no-unbounded-retry-growth
quality:
- task-and-grounding-gates-pass
- tool-selection-gate-passes
control:
- refusal-parity-passes
- trace-completeness-passes
- no-duplicate-tool-operation
automatic_stop:
- write-dispatched-without-valid-approval
- state-ownership-conflict
- missing-region-or-version-in-trace
- retrieval-snapshot-mismatch Set thresholds from the service’s observed baseline and risk policy rather than copying generic percentages. One control breach should stop the canary even when average latency improves.
Correlate the failover in traces
Microsoft Foundry tracing can feed agent execution telemetry into Application Insights. Whatever instrumentation path is used, the failover investigation needs a continuous chain across request, model, retrieval and tools.
let IncidentStart = datetime(2026-09-04T07:42:00Z);
let IncidentEnd = IncidentStart + 45m;
AgentRunEvents
| where TimeGenerated between (IncidentStart .. IncidentEnd)
| where AgentName == "operations-assistant"
| summarize
Runs = dcount(RunId),
FailedRuns = dcountif(RunId, RunStatus != "completed"),
P95LatencyMs = percentile(DurationMs, 95),
ToolCalls = countif(EventType == "tool_call"),
DuplicateOperations = dcountif(OperationId, IsDuplicate == true),
MissingTraceContext = countif(isempty(TraceId) or isempty(DeploymentVersion))
by Region, DeploymentName, AgentVersion, bin(TimeGenerated, 5m)
| extend FailureRate = todouble(FailedRuns) / Runs
| order by TimeGenerated asc Adapt the table and field names to the telemetry contract. Pair operational metrics with sampled evaluations: low error rates do not reveal grounding drift, changed tool choices or weakened refusals.
Decide hold, fail over, fail back or rollback
The runbook ends with an explicit traffic and control decision.
Hold on the primary path
Primary degradation is not reproduced or dependency scope is unknown
Candidate configuration, state or traces cannot be proved
Recovery risk is greater than the bounded incident impact
Fail over read-only traffic
Candidate probes and evaluation gates pass
New sessions can be isolated and traced
Write tools remain blocked
Promote the candidate path
Canary meets operational, quality and control gates
State ownership and idempotency are proven
Tool, identity and approval boundaries match the approved contract
Fail back
Primary health is stable through the observation window
New sessions return gradually with session affinity
Existing candidate runs finish or are explicitly reconciled
Rollback the routing change
Quality, refusal, trace or state gate fails
Candidate weights return to zero
Ambiguous actions are reconciled before any retry
Incident evidence retains both regional paths Automate probe collection, traffic weights and stop conditions where their semantics are deterministic. Keep the production-write decision behind an explicit gate until the secondary path has proved the same action boundary as the primary.
Conclusion
Regional failover for an AI agent is a change to an operational system, not a DNS shortcut around an unhealthy model. The model may recover while state, retrieval, tools, identities or approval controls diverge.
Fail over the smallest useful workload, prove versions and policies, evaluate outcomes, protect side effects, and canary with trace continuity. Promote only when availability, quality and control pass together. Otherwise hold or return traffic, reconcile every ambiguous action, and keep the rollback narrower than the incident.