AI

Microsoft Foundry Toolbox: validate a version before promoting it to agents

A production runbook to test a Toolbox version, its MCP tools, identities, approvals and traces before promotion, then return to the previous version without redeploying agents.

19 Aug 2026 microsoft-foundryaiagentstoolboxmcpagentopsidentityapprovalsobservabilityevaluationguardrailsrunbookrollbackproduction

A platform team centralizes Azure AI Search, an internal MCP server and an OpenAPI operation in a Microsoft Foundry Toolbox. Several agents consume the same endpoint. A new version adds an operation, changes an input schema and tightens authentication. Promoting it as the default may look safer than redeploying applications, yet the change can reach every agent that follows that version without a code change.

The production case is an operations agent that searches runbooks, reads service state and prepares an action for approval. The goal is not merely to see tools in tools/list. The team must prove the contract, effective identity, approval boundary, traces and failure behavior before moving a shared pointer. Rollback means restoring the previous default version, not assembling another configuration during the incident.

Identify every consumer of the default version

A Toolbox is a named, versioned bundle of tool configurations exposed through an MCP-compatible endpoint. Centralization removes duplication, but it also turns promotion into a platform change. Start by inventorying consumers and how they resolve a version.

yaml toolbox-consumer-inventory.yml
toolbox:
name: ops-toolbox
current_default: v17
candidate: v18
rollback_target: v17

consumers:
- agent: incident-triage
  resolution: default-version
  runtime_identity: mi-agent-triage-prod
  sensitive_tools: [restart_job]
- agent: runbook-search
  resolution: pinned-v17
  runtime_identity: mi-agent-search-prod
  sensitive_tools: []
- client: vscode-operations
  resolution: default-version
  user_identity_passthrough: true

change_scope:
added_tool: get_change_window
changed_schema: restart_job
changed_connection: operations-api
forbidden_drift: [search_index, approval_policy]

A pinned consumer does not receive the release at the same time as one following the default. The inventory must also separate the developer identity that manages the Toolbox, the agent managed identity used at runtime, and any end-user identity passed through an OAuth flow.

Freeze the candidate contract

Create a candidate version without promoting it. The change record keeps the definition, connections, guardrail policy and a digest of the client-visible contract. Never use latest as evidence.

json toolbox-release-manifest.json
{
"toolbox": "ops-toolbox",
"version": "v18",
"previousDefault": "v17",
"endpointMode": "version-pinned",
"toolContractDigest": "<sha256>",
"connections": [
  {"name": "runbook-search", "auth": "managed-identity"},
  {"name": "operations-api", "auth": "oauth-or-managed-identity"}
],
"approvalPolicy": "approvals-v6",
"guardrailPolicy": "toolbox-rai-v4",
"owner": "platform-ai",
"expiresIfNotPromoted": "2026-08-26T18:00:00Z"
}

Calculate the digest from a normalized representation of names, descriptions, input schemas and control metadata. It must contain no secret or token. If a connection or policy changes between testing and promotion, invalidate the evidence and qualify the new candidate again.

Test the versioned endpoint as an MCP client

Retrieve the candidate endpoint and connect an isolated client. The minimum sequence is initialize, followed by tools/list. Confirm that the list is not empty, every tool has a name, useful description and usable inputSchema, and names remain stable and correctly namespaced.

text toolbox-mcp-contract-checks.txt
Initialization
HTTP 200 and MCP session initialized
endpoint explicitly bound to v18
no implicit redirect to the default version

tools/list
non-empty list
unique and stable names
description precise enough for tool selection
inputSchema.properties present
required fields, types and enums compatible with v17
removed tool absent only when the break is approved

Negative controls
unknown argument rejected
out-of-scope production target rejected by the server
identity without the role denied
sensitive call without approval not executed
timeout and dependency failure returned as errors

Compare v17 and v18 structurally. A description change can alter model selection even when the JSON Schema is identical. A broader enum, a newly optional field or a server-side default can widen the real action without breaking the client.

Prove identities and approval enforcement

Use the identities that will exist in production, against tightly bounded resources. Success with an administrator does not validate the agent’s managed identity. A 403 must remain an authorization failure, not become an empty result the model can interpret as valid state.

The require_approval metadata published with a tool tells the runtime that confirmation is required. By itself, it is not a server-side barrier: the consuming runtime must display the pending action, wait for a decision and prevent the call until approval is granted. Sensitive operations also need independent server controls such as a limited role, resource scope, idempotency key and strict parameter validation.

yaml toolbox-sensitive-tool-policy.yml
tool: operations.restart_job
require_approval: always
runtime_enforcement:
show: [environment, job_id, reason, correlation_id]
approver_role: production-operator
approval_ttl_seconds: 300
server_enforcement:
allowed_environments: [preprod, production]
production_scope: [job-billing-nightly]
require_idempotency_key: true
reject_unknown_arguments: true
max_execution_seconds: 30
evidence:
log_arguments: redacted
log_result: status-only
retain_approval_id: true

Exercise denial, expiry and attempted reuse. A UI that shows confirmation without binding the decision to the exact arguments can let a different call inherit stale approval.

Evaluate scenarios, not just connectivity

A successful ping does not show whether the agent selects the right tool. Replay a bounded set of representative and adversarial requests against the candidate: documentation search, state lookup, an allowed action, an out-of-scope action, ambiguous input, a slow dependency and a malformed tool response.

yaml toolbox-promotion-gates.yml
blocking:
- forbidden_tool_selected
- approval_bypassed
- broader_resource_scope
- secret_or_token_in_trace
- tool_error_reported_as_success
- write_retried_without_idempotency

measured:
- tool_call_accuracy
- task_adherence
- correct_refusal
- argument_schema_validity
- p95_tool_latency
- trace_completeness

decision:
blocking_failures_allowed: 0
sensitive_cases_require_human_review: true
aggregate_score_cannot_override_blocking_failure: true

Keep the prompt, model and test set fixed during the comparison. Otherwise, an improvement or regression cannot be attributed to the Toolbox release.

Verify traces without leaking arguments

Traces should connect the agent run, Toolbox name and version, selected tool, approval decision, latency and final status. Arguments and results may contain sensitive data, so redact them before ingestion and retain only what incident investigation needs.

kusto toolbox-candidate-evidence.kql
AgentToolInvocations
| where TimeGenerated > ago(2h)
| where ToolboxName == "ops-toolbox" and ToolboxVersion == "v18"
| summarize
  Calls = count(),
  Failures = countif(Status != "success"),
  ApprovalBypass = countif(RequiresApproval and ApprovalDecision != "approved" and Executed),
  MissingTrace = countif(isempty(AgentTraceId) or isempty(ToolCallId)),
  P95LatencyMs = percentile(DurationMs, 95)
by ToolName, RuntimeIdentity
| order by ApprovalBypass desc, Failures desc

AgentToolInvocations is a normalized example table to adapt to the actual Application Insights or OpenTelemetry schema. Inspect runtime telemetry too: a healthy Toolbox with missing traces is still not operable in production.

Promote in stages and keep the return path ready

Promotion is allowed when the contract is frozen, negative controls pass, least-privileged identities work, no approval is bypassed, blocking evaluations are at zero and traces are complete. Start with a validation agent pinned to v18, then a canary of explicit consumers before changing the default.

At change time, record old and new versions, UTC timestamp, validated digest and expected consumers. Watch tool selection, 401/403, schema errors, timeouts, approval requests and business effects. A lower call count can mean that tools are no longer discovered; it is not automatically an improvement.

Roll back to v17 when a tool disappears, scope broadens, an identity fails, approval is no longer bound to arguments or traces become incomplete. Restore the previous default pointer, verify its digest and replay critical checks. Agents pinned to v18 need separate treatment because changing the default does not move them back.

Conclusion

A Toolbox centralizes tools, connections and controls, but it also centralizes change risk. The production unit is therefore not an endpoint that responds. It is an immutable version with a proven contract, identities, approvals, evaluations and traces.

The final decision is deliberate: promote v18 through a canary with a tested return path, keep v17 while evidence is incomplete, or roll back immediately when a sensitive invariant fails. This keeps shared agent tooling evolvable without turning a central promotion into a silent production change.