AI

AgentOps: prevent MCP token forwarding before production actions

A production runbook for proving token audience, delegation, scopes and runtime identity across an MCP tool chain before restoring state-changing actions.

13 Sept 2026 aiagentopsmcpoauthentra-ididentitysecurityobservabilityguardrailsautomationrunbookrollbackproduction

An internal operations agent can read incidents correctly and still act under the wrong authority. A user asks it to inspect one service, the agent calls an MCP server, and the downstream API records a broad application identity or a token intended for another resource. The request works, but the authorization chain no longer proves who asked, what was approved, or why the tool could reach that target.

The running case is an incident assistant connected to Microsoft Foundry. It calls an internal MCP server to read service health and prepare a bounded restart. Read operations use the caller’s delegated context, while state changes require a dedicated workload identity and human approval. After an authentication change, traces no longer agree on the token audience and acting principal. This runbook leads to one decision: restore the intended delegation, keep writes disabled behind read-only tools, or roll back the authentication release.

Freeze one authorization chain

Do not begin by granting another role. Capture one request from the user boundary to the target API and record identifiers, not bearer tokens. A complete chain distinguishes the human, the agent runtime, the MCP server and the downstream resource.

yaml mcp-authorization-incident.yml
incident: inc-20260913-006
request:
intent_id: intent-7f21
trace_id: 8cf4b7d6a2d14ef1
user_object_id: <user-oid>
approved_action: prepare_restart
approved_target: api-orders-prod / instance-03
agent:
runtime: ops-assistant-prod
release: 2026.09.13.1
mcp:
server: mcp-operations-prod
tool: prepare_service_restart
tool_call_id: call-31c8
downstream:
api: operations-control-api
expected_audience: api://operations-control
expected_mode: workload-identity-after-human-approval
preserve:
- token_claims_redacted
- approval_record_and_fingerprint
- tool_arguments_and_schema_version
- downstream_request_id
- identity_and_policy_release_timeline

Never place access tokens, authorization headers or refresh tokens in the incident record. Keep only the minimum claims needed to explain the decision, with access restricted to the investigation.

Name the authority at every hop

“The agent is authorized” is too vague. Each hop needs a subject, audience, credential mode and allowed action. Delegated access and application access answer different questions: the first carries a user context; the second represents a workload acting under its own permissions.

text mcp-authority-map.txt
Hop 1  User -> agent runtime
Subject: user object ID and tenant
Audience: agent application
Purpose: submit an operational request

Hop 2  Agent runtime -> MCP server
Subject: delegated user or agent workload identity
Audience: MCP server
Purpose: invoke one declared tool

Hop 3  MCP server -> downstream API
Subject: explicit delegated exchange or dedicated MCP workload identity
Audience: downstream API
Purpose: read health or prepare one bounded action

Forbidden shortcuts
Forward the original bearer token without validating its audience
Reuse one broad application token for unrelated downstream APIs
Treat tool visibility as permission to execute it
Convert a user request into an application write without approval

This map is the expected contract. The incident is the difference between this contract and what the runtime actually emitted.

Inspect claims without exposing credentials

Inspect a token only in a controlled diagnostic path. Validate its signature and policy in the receiving service; local decoding is useful for triage but is not proof of authenticity. Log a one-way token fingerprint and a small allowlist of claims so the same credential can be correlated across hops without becoming replayable evidence.

javascript redact-token-claims.mjs
import { createHash } from 'node:crypto';

export function diagnosticEnvelope(rawToken, verifiedClaims) {
return {
  tokenFingerprint: createHash('sha256').update(rawToken).digest('hex').slice(0, 16),
  iss: verifiedClaims.iss,
  aud: verifiedClaims.aud,
  tid: verifiedClaims.tid,
  oid: verifiedClaims.oid,
  azp: verifiedClaims.azp ?? verifiedClaims.appid,
  scp: verifiedClaims.scp?.split(' ') ?? [],
  roles: verifiedClaims.roles ?? [],
  iat: verifiedClaims.iat,
  exp: verifiedClaims.exp
};
}

At each receiver, verify issuer, tenant, audience, lifetime and the expected permission representation. A delegated token commonly exposes scopes, while an application token commonly exposes application roles. Do not accept either shape merely because the claim exists; compare it with the endpoint’s explicit policy.

Correlate the caller seen by the target

Agent traces prove what the orchestrator attempted. The downstream API proves which identity arrived. Join both views with trace_id, tool_call_id and the downstream request ID. If the target sees an unexpected application identity, expanding that identity’s role would preserve the defect.

kusto mcp-token-chain-correlation.kql
let StartTime = datetime(2026-09-13T06:00:00Z);
let EndTime = datetime(2026-09-13T07:00:00Z);
let TraceId = "8cf4b7d6a2d14ef1";
union isfuzzy=true
(AppTraces
 | where TimeGenerated between (StartTime .. EndTime)
 | where tostring(Properties.trace_id) == TraceId
 | project TimeGenerated, Source="agent-or-mcp", Message,
           Tool=tostring(Properties.tool),
           ToolCallId=tostring(Properties.tool_call_id),
           Audience=tostring(Properties.token_aud),
           Caller=tostring(Properties.caller_oid),
           ClientApp=tostring(Properties.client_app_id),
           Result=tostring(Properties.result)),
(AppRequests
 | where TimeGenerated between (StartTime .. EndTime)
 | where tostring(Properties.trace_id) == TraceId
 | project TimeGenerated, Source="downstream-api", Name,
           Tool="", ToolCallId=tostring(Properties.tool_call_id),
           Audience=tostring(Properties.token_aud),
           Caller=tostring(Properties.caller_oid),
           ClientApp=tostring(Properties.client_app_id),
           Result=tostring(ResultCode))
| order by TimeGenerated asc

Property names depend on the telemetry contract. Define them at the authentication middleware and tool wrapper instead of parsing free-text logs. Absence of the downstream line is also meaningful: the call may have failed before authentication, been blocked by policy, or followed another network path.

Separate delegation from the privileged action

A useful pattern is to keep diagnosis and execution distinct. Read-only tools may use a bounded delegated identity when the target API supports it. A state-changing tool should require an approval record, an exact target and a dedicated workload identity whose permissions match the operation. The MCP server must not silently promote one mode into the other.

json mcp-write-authorization-contract.json
{
"tool": "prepare_service_restart",
"schemaVersion": "3.2",
"mode": "prepare_only",
"required": {
  "intentId": "intent-7f21",
  "target": "api-orders-prod/instance-03",
  "expectedState": "unhealthy",
  "approvalFingerprint": "sha256:<fingerprint>",
  "expiresAt": "2026-09-13T07:15:00Z"
},
"identityPolicy": {
  "acceptedAudience": "api://operations-control",
  "acceptedClient": "<mcp-workload-client-id>",
  "requiredApplicationRole": "ServiceRestart.Prepare",
  "delegatedScopesAccepted": false
},
"rejectWhen": [
  "target_differs_from_approval",
  "approval_expired",
  "caller_or_audience_unexpected",
  "execution_requested_in_prepare_only_mode"
]
}

Tool schemas constrain arguments; API authorization constrains effects. Both are required. A hidden tool is not a disabled permission, and a valid token is not approval for every parameter.

Contain the narrowest boundary

If the chain is ambiguous, remove write capability without disabling diagnosis. Pin the previous authentication configuration when it is known healthy, or force sensitive tools into prepare_only. Revoke the affected credential only when you know what else depends on it; blind revocation can hide the original caller and break unrelated reads.

text mcp-token-containment.txt
Immediate containment
Disable execute mode for state-changing MCP tools
Keep approved source retrieval and read-only health checks
Reject tokens whose issuer, tenant or audience is not exact
Reject delegated tokens on application-only endpoints
Preserve one failing and one healthy trace

Do not do during triage
Add a broad API role to the identity currently observed
Log raw bearer tokens for easier comparison
Accept multiple audiences temporarily
Re-enable writes because the same request succeeds manually
Rotate every credential before mapping consumers

The objective is to stop unjustified authority while keeping enough observability to repair the contract.

Exercise positive and negative authorization paths

Replay with synthetic identities and non-production targets first. Test the intended path and deliberate violations: wrong audience, wrong tenant, missing scope or role, expired approval, changed target, reused approval fingerprint and an attempt to pass a delegated token to an application-only operation.

yaml mcp-authorization-evaluation.yml
cases:
- id: delegated-health-read
  token: delegated / Health.Read / correct audience
  tool: get_service_health
  expect: allowed
- id: wrong-audience
  token: valid signature / different audience
  tool: get_service_health
  expect: denied_before_tool_execution
- id: delegated-write
  token: delegated / correct audience
  tool: execute_service_restart
  expect: denied_application_role_required
- id: stale-approval
  token: workload / ServiceRestart.Execute
  approval: expired
  expect: denied_before_downstream_call
- id: target-substitution
  token: workload / ServiceRestart.Execute
  approved_target: instance-03
  requested_target: all-instances
  expect: denied_target_mismatch
promotion_gate:
unexpected_callers: 0
raw_tokens_in_logs: 0
negative_cases_allowed: 0
trace_chain_complete: true

Then canary one read path and one prepared action. Real execution comes last, on one reversible target, with the target API recording the expected client, role, approval fingerprint and correlation IDs.

Decide restore, degrade or roll back

Restore state-changing tools only when every hop accepts the intended audience and identity, the target enforces the right scope or role, approval is bound to the exact action, and negative tests fail closed. Keep the agent in read-only or prepare_only mode when the authorization chain remains useful but cannot yet prove execution authority.

Roll back the authentication release when the previous contract is known, compatible and narrower. Rotate or revoke a credential when it was exposed or used outside policy, after mapping its consumers. If the design depends on passing one token through several services without an explicit delegation contract, stop the rollout and redesign that hop rather than documenting the ambiguity as an exception.

Conclusion

An MCP tool call is not one authorization event. It is a chain of receivers, audiences, subjects, policies and effects. Diagnose it by comparing the expected authority map with verified claims and the identity observed by the downstream API.

The closing decision is operational: restore explicit delegation, use a dedicated workload identity for approved writes, keep sensitive tools in a non-executing mode, or roll back the release. The incident is closed only when the intended call succeeds, substituted authority is rejected, raw tokens are absent from logs and the full action remains attributable.