AI
AgentOps: contain sensitive data in agent traces before disabling observability
A production runbook for bounding a leak in agent traces, locating the unsafe field, canarying redaction and restoring useful observability.
After an operations agent release, a team finds a span containing the full body of a ticket: a name, an email address, a log excerpt and a temporary token pasted by mistake. Disabling every trace may limit further exposure, but it also removes the evidence needed to identify affected sessions, tools and export routes.
The running case is a Microsoft Foundry agent that reads runbooks, inspects tickets and prepares bounded actions. An instrumentation change added tool inputs and outputs to trace attributes. This runbook leads to four decisions: isolate the unsafe path, handle data already exported, restore clean minimum telemetry, then promote corrected fields or roll back the instrumentation.
Freeze the incident without copying the data
Start with the trace identifier, release, UTC time, span type, tool, exporter and destination. Do not paste the sensitive value into the incident ticket. Retain a fingerprint, a data class and a pointer to restricted evidence.
incident:
first_seen_utc: 2026-10-04T15:42:18Z
agent_release: ops-agent-2026-10-04.3
trace_id: <trace-id>
span_name: tool.ticket.read
tool: ticket.read
exporter_route: agent-traces-prod
destinations: [primary_observability, security_archive]
suspected_field:
attribute: tool.result.body
data_class: personal_and_secret
fingerprint_sha256: <restricted-fingerprint>
raw_evidence: <restricted-vault-reference>
containment_owner: platform-security
decision_deadline_utc: 2026-10-04T17:00:00Z The fingerprint supports matching without redistributing the content. If the value is a secret, treat it as compromised once it reaches a readable destination and start revocation in parallel with the investigation.
Narrow the flow to the smallest boundary
Identify where the value becomes telemetry. It may enter through a user message, a retrieved document, a tool result, an exception, an SDK attribute or an enrichment processor. It may then be copied to several destinations.
Possible inputs
user message | retrieval document | tool result | exception
Copy points
tool wrapper -> agent instrumentation -> collector -> exporter -> destination
Decision at each point
Keep: IDs, versions, status, latency, size, policy decision
Transform: business IDs to stable hashes when correlation is required
Drop: raw prompts, tool bodies, headers, tokens, secrets, attachments
Restrict: exceptional raw evidence with bounded access and retention Useful containment does not shut down the whole agent when one tool wrapper is faulty. Disable the affected enrichment, temporarily remove the tool from the catalog or force it into read-only mode, while keeping minimum events: trace ID, tool, runtime identity, policy decision, status and error code.
Stopping only the final exporter is insufficient when a queue, collector, secondary store or debug log already holds the payload. Map every copy before declaring the path contained.
Measure exposure through metadata
Query attribute names, versions and export routes before searching for the sensitive value in clear text. The query below assumes a normalized AgentTraceEvents table; adapt its columns to the actual pipeline.
let Start = datetime(2026-10-04T14:00:00Z);
let End = datetime(2026-10-04T16:30:00Z);
let SuspectRelease = "ops-agent-2026-10-04.3";
AgentTraceEvents
| where TimeGenerated between (Start .. End)
| where AgentRelease == SuspectRelease
| where AttributeNames has_any ("tool.result.body", "prompt.content", "http.request.header.authorization")
| summarize
Events = count(),
FirstSeen = min(TimeGenerated),
LastSeen = max(TimeGenerated),
Traces = dcount(TraceId),
Conversations = dcount(ConversationId)
by SpanName, ToolName, ExporterRoute, Destination
| order by Events desc Search for the fingerprint only inside an authorized function or environment, returning identifiers rather than payloads. Event volume is not the affected-person count: one value may be repeated across several spans and destinations.
Define a positive telemetry contract
A denylist of patterns can never cover every sensitive value. Define the fields allowed for each event type instead. Any undeclared attribute must be dropped, truncated or routed to restricted storage under an explicit rule.
event: agent.tool.completed
allowed:
- trace_id
- conversation_id_hash
- agent_release
- tool_name
- tool_contract_version
- policy_decision
- status_code
- duration_ms
- input_bytes
- output_bytes
forbidden:
- raw_prompt
- tool_arguments_raw
- tool_result_body
- authorization_header
- retrieved_document_content
limits:
string_length: 256
attribute_count: 32
restricted_debug:
enabled: false
approval_required: security_incident_owner
max_retention_hours: 4 Enforce this contract as close to the source as possible, then again before export. The first boundary avoids spreading data through the pipeline; the second protects against instrumentation or a service that bypasses the expected wrapper.
Canary redaction with synthetic markers
Do not validate the fix with a real address, token or customer excerpt. Inject unique, non-secret markers and classify them as though they were sensitive. Send them separately through the message, retrieval, tool arguments, tool result and exception paths.
canaries:
- id: user_input_marker
path: user.message
value: NAXAYA_CANARY_USER_20261004
- id: retrieval_marker
path: retrieval.document.content
value: NAXAYA_CANARY_RETRIEVAL_20261004
- id: tool_result_marker
path: tool.result.body
value: NAXAYA_CANARY_TOOL_20261004
- id: exception_marker
path: exception.message
value: NAXAYA_CANARY_ERROR_20261004
expected:
allowed_destinations: []
metadata_preserved: [trace_id, tool_name, status_code, policy_decision]
terminal_action: read_only The test passes only when no raw marker reaches ordinary destinations, expected metadata remains correlated and failures remain diagnosable. Redaction that turns every value into an empty string can make the pipeline compliant but unusable.
Handle copies already exported
Fixing the pipeline does not erase history. For every destination, identify its owner, readers, secondary exports, backups, retention and deletion or restriction capabilities. Retain the timeline and evidence IDs even when the payload must be removed.
Revoke an exposed secret; do not rely on log deletion alone. For personal or business data, follow the organization’s security and privacy process. The technical runbook should produce the list of traces, destinations and potential access, not decide notification obligations on its own.
Restore progressively and retain rollback
Deploy the telemetry policy to a canary cohort. Restore read-only diagnostics first, then tools that do not return free-form content. Write-capable tools return only after their arguments, results, refusals and error paths pass validation.
Promote
No raw marker in ordinary destinations
Allowed fields present and correlated end to end
Errors and refusals remain usable
Revocation and historical handling completed or tracked
Keep containment
A secondary route or cache is not qualified
The contract does not cover one span type
Destination access or retention remains unknown
Roll back instrumentation
Redaction breaks correlation or hides critical failures
Restore the previous known-safe bundle
Keep the faulty wrapper and sensitive tools disabled
Replay every canary before another promotion A rollback must never restore the raw field that caused the incident. It returns to the last safe instrumentation while the faulty path stays isolated. Close the incident only when the current flow is clean, historical data is handled, exposed secrets are revoked and canaries prove both confidentiality and operability.
Conclusion
Sensitive data in an agent trace is a pipeline incident, not sufficient reason to become blind. The useful response ties release, span, field, route and destination together; it preserves the metadata needed for operations while isolating raw content.
The final decision is explicit: promote canary-tested redaction, keep containment while any copy remains unknown, or restore a safe instrumentation bundle. The goal is not fewer traces. It is traces that explain actions without becoming a leak themselves.