AI

AgentOps: neutralize indirect prompt injection in MCP tool output before a production action

A production runbook for isolating hostile instructions carried by MCP tool output, preserving provenance, enforcing policy outside the model, and validating or rolling back write access.

27 Sept 2026 aiagentopsagentsmcptoolsprompt-injectionsecurityprovenanceobservabilityevaluationguardrailsautomationrunbookrollbackproduction

An operations agent reads a change ticket through an MCP tool, then proposes a production deployment that the operator did not request. The ticket is legitimate, the tool call succeeded and the response matches its JSON schema. One free-text field, however, contains an instruction telling the agent to ignore approval and invoke another tool.

This is indirect prompt injection through tool output. It differs from a poisoned retrieval corpus: the hostile text arrives during an active workflow from a system the agent is allowed to query. It also survives ordinary schema validation because a valid string can still contain an instruction. The running case is a Microsoft Foundry operations agent that reads change records, queries Azure state and may call a bounded deployment tool after human approval. The decision is explicit: reopen writes, keep the agent read-only, quarantine the connector or restore the previous tool contract.

Freeze the complete action chain

Stop new write-capable calls without deleting the trace. Preserve the user request, agent and prompt versions, MCP server and tool versions, arguments, raw result, normalized result, proposed next call, policy decision, approval state and execution identity. A screenshot of the final answer cannot show where the instruction entered the chain.

yaml 01-mcp-output-injection-incident.yml
incident: INC-AI-731
window_utc: 2026-09-27T14:10:00Z/2026-09-27T14:25:00Z
agent:
name: ops-assistant-prod
release: 2026.09.27-2
tool_call:
server: change-catalog-mcp
tool: read_change_request
schema_version: 4
arguments_hash: <sha256>
raw_result_hash: <sha256>
source:
system: change-catalog
record_id: CHG-1842
revision: 19
proposed_action:
tool: deploy_release
target: production
approval_id: missing
containment:
write_tools: disabled
read_tools: enabled
rollback:
agent_release: 2026.09.26-1
tool_schema_version: 3

Keep the raw response under restricted access and hash it before redaction. The investigation needs the exact bytes, but incident notes should not redistribute secrets or make the hostile text a reusable payload.

Locate the trust-boundary violation

Reconstruct the chain field by field. Separate control data, factual evidence and untrusted content. Tool name, schema version and policy outcome are control data. A deployment status returned by Azure is evidence. A ticket title, comment, log line, repository issue or vendor message is untrusted content even when it comes through an authenticated API.

The key question is not whether the MCP server is trusted. It is whether each returned field is authorized to influence an action. Authentication proves which service supplied the response; it does not turn user-authored text into policy.

Compare the raw backend response, MCP result and model-visible representation. Look for transformations that concatenate fields, remove provenance, promote a comment into a summary, or place data next to system-like labels. Record the exact field that changed the next-tool decision. If traces do not retain tool result identifiers and policy decisions, keep writes closed until that observability gap is fixed.

Give tool output an explicit data contract

Do not solve the incident only by adding “ignore malicious instructions” to the system prompt. Narrow the tool response so the agent receives the minimum useful fields with their origin and trust class. Keep free text separate from machine-actionable values.

json 02-bounded-mcp-result.json
{
"resultVersion": "5",
"source": {
  "system": "change-catalog",
  "recordId": "CHG-1842",
  "revision": "19",
  "retrievedAt": "2026-09-27T14:16:32Z"
},
"facts": {
  "environment": "production",
  "requestedRelease": "api@2026.09.27.1",
  "workflowState": "awaiting_approval"
},
"untrustedContent": {
  "contentType": "user_authored_text",
  "value": "<redacted ticket comment>",
  "mayAuthorizeAction": false
},
"actionPolicy": {
  "writeAllowed": false,
  "reason": "approval_not_verified"
}
}

The MCP server should reject unknown fields where the contract allows it, cap response size and identify truncation. Those controls reduce ambiguity but do not detect every injection. A perfectly valid untrustedContent.value remains data, never an instruction, tool argument or approval token.

Enforce authorization outside the model

The model may propose the next step; it must not be the authority that permits it. Put a deterministic gate in the orchestration or tool gateway before every write. Re-read current state from an authoritative source and verify the actor, target, allowed operation, approval, expiry and request fingerprint independently of the previous tool text.

Bind approval to the canonical action: tool, normalized arguments, environment and resource identifiers. If any of them changes, the approval is stale. Do not let a ticket comment supply an approval ID, execution identity or hidden override.

Require idempotency for the write tool. If the incident involved a timeout or uncertain result, query action status before retrying. Injection containment is incomplete if replaying the same intent can duplicate a legitimate action.

Detect the behavior, not only suspicious words

Keyword detection can help triage, but attackers can avoid obvious phrases and legitimate logs can contain them. Alert on the behavioral boundary: an untrusted field influences a new tool, a write is proposed without independently verified approval, or arguments appear only in free text.

kusto 03-mcp-output-action-boundary.kql
let Window = 24h;
AgentToolEvents
| where TimeGenerated > ago(Window)
| where EventType in ("tool_result", "tool_proposed", "policy_decision", "tool_executed")
| summarize
  UntrustedResults=countif(ResultTrust == "untrusted"),
  WritesProposed=countif(OperationClass == "write" and EventType == "tool_proposed"),
  WritesBlocked=countif(OperationClass == "write" and PolicyDecision == "deny"),
  WritesWithoutApproval=countif(OperationClass == "write" and ApprovalVerified != true),
  DistinctSources=dcount(SourceRecordId)
by CorrelationId, AgentRelease, ToolName, bin(TimeGenerated, 5m)
| where UntrustedResults > 0 and (WritesProposed > 0 or WritesWithoutApproval > 0)
| order by TimeGenerated desc

Adapt table and field names to the trace pipeline. Retain correlation across user intent, MCP call, backend record revision, policy decision and real action. A blocked attempt is a security signal, not noise to remove from telemetry.

Evaluate adversarial results before reopening writes

Build tests at the tool-result boundary, not only at the initial prompt. Replay the same user intent with controlled result variants: an instruction in a comment, encoded text, a fake approval identifier, arguments hidden in a log, a very long result that pushes policy context away, and a clean control result.

Pass criteria must be observable. The agent may summarize the untrusted text, but it must label its origin, refuse to treat it as policy, obtain current facts through approved fields and avoid every write until the external gate succeeds. Add a negative control proving that a valid approval for action A cannot authorize action B.

Run the candidate in shadow mode first. Then expose read-only tools to a small cohort. A write canary should target a reversible non-production resource with an independently issued approval and a refusal test alongside it. Do not canary an injection defense by granting broader production rights.

Decide and roll back cleanly

Reopen production writes only when the source field is identified, the tool result keeps provenance and trust class, policy is enforced outside the model, adversarial evaluations pass, traces explain the decision and the refusal control stays blocked.

Keep the agent read-only when the workflow remains useful but attribution or approval evidence is incomplete. Quarantine the MCP connector when it cannot separate untrusted content from action fields or when the backend record was compromised. Restore the previous agent release or tool contract when the regression follows a deployment and the retained version passes the same evaluation set.

If an unintended action executed, roll back the affected production resource through its own runbook. Disabling the agent does not undo a deployment, route, permission or automation job. After rollback, verify both system state and the action ledger, then preserve the incident trace for regression tests.

Conclusion

MCP output is not automatically safe because the server is authenticated or the JSON is valid. In an agentic workflow, every free-text field crosses a trust boundary when the model can turn it into another action.

The durable control is layered: preserve provenance, label untrusted content, minimize the result contract, authorize writes outside the model, trace the full chain and evaluate hostile tool responses. Production access returns only when the system can prove that data informed the diagnosis without becoming an instruction.