AI
AgentOps: validate session and memory isolation before production
A production runbook for detecting context leaks across AI agent sessions, separating memory, caches, retrieval, tools and identity, then enabling or rolling back persistence without losing audit traces.
An operations agent supports two teams working on unrelated incidents. In a fresh conversation, it mentions a service name, ticket identifier or tool result that only appeared in the previous session. The answer may look relevant, but the failure is severe: a session, user or environment boundary is no longer isolated.
This runbook covers an internal agent that reads runbooks, queries read-only signals and may keep working or durable memory. Whether the implementation uses Microsoft Foundry or another runtime, the production question is the same: before persistence reaches more users, can the team prove that every memory, retrieved document and tool result remains bound to its authorized scope?
Define the isolation contract before testing
“Memory” often hides several states: conversation history, compacted summary, durable profile, retrieval cache and tool output. Name the boundaries each layer must enforce before running a test.
agent: ops-assistant-prod
boundaries:
user: required
team: required
environment: required
session: required
state_layers:
conversation_history:
scope: session
compacted_summary:
scope: session
durable_memory:
scope: user_and_team
write_policy: approved_facts_only
retrieval_cache:
scope: corpus_version_and_access_scope
tool_results:
scope: session
forbidden:
- reuse another session tool result
- expose another team incident marker
- retrieve a document outside caller scope
- persist secrets or raw diagnostic payloads
release_decision:
- enable
- keep_read_only
- disable_persistence
- rollback_runtime_version A session must not be the only boundary. If its identifier is predictable, reused or accepted without checking user and logical tenant, changing session_id is not enough. Authorization must be recomputed for every memory read, retrieval operation and tool call.
Seed synthetic canaries in every scope
An isolation test must actively look for leakage, not just confirm that the expected answer appears. Create non-secret, non-customer markers that differ by user, team, environment and session.
actors:
- user: alice.ops
team: payments
environment: production
session: session-payments-a
marker: CANARY-PAYMENTS-7K4Q
allowed_source: runbook-payments-v12
- user: bob.ops
team: logistics
environment: production
session: session-logistics-b
marker: CANARY-LOGISTICS-9M2R
allowed_source: runbook-logistics-v8
tests:
- create each session independently
- write its marker only through an approved path
- ask the other session for recent incident context
- call the same read-only tool with different scopes
- expire and reopen one session
- repeat requests concurrently
pass:
- expected marker appears only in its owning scope
- foreign marker is absent from answers, citations and tool arguments
- inaccessible source is not retrieved
- denial is explicit and traced Do not seed real confidential data to make the exercise realistic. The canary exists to reveal a boundary crossing. It should be searchable in traces without creating a security incident of its own.
Isolate each state layer
If a marker appears in the wrong scope, disabling all memory hides the source. Replay the same journey while enabling one state layer at a time.
Start without persistence: a new session, no imported summary, a cold cache and read-only tools. Then enable history, compaction, durable memory, cached retrieval and tools in that order. The first layer that makes the foreign marker reappear identifies the failure family.
The marker leaks without a tool call
Inspect history loading, summaries and the session key
The marker leaks after compaction
Inspect summary scope and storage
The marker leaks only after a new login
Inspect durable memory, profiles and user resolution
A foreign document appears in citations
Inspect access filters before search and the retrieval cache
The marker appears in tool arguments or output
Inspect identity binding, tool caches and response reuse
The leak appears only under concurrency
Inspect global state, worker pools and context propagation This sequence keeps prompt behavior separate from architecture. An instruction saying “do not use another user’s data” cannot repair an incomplete cache key or shared worker state.
Exercise concurrency and worker reuse
Sequential tests often pass while production fails under load. Run both sessions in parallel, interleave requests and force reuse of the same worker pool where the architecture allows it. Include streaming responses, retries and timeouts: a late callback can write into whichever session now occupies that worker.
Inspect actual keys, not their labels in source code. A retrieval cache key needs corpus identity or version and access scope. Durable memory needs a logical owner. A tool result must not be reused merely because another session issued similar text.
Stop the test when
A foreign canary appears in an answer or citation
A tool receives another actor's scope
A trace cannot link user, session and execution
A retry continues after session expiry
Test cleanup removes audit evidence
Preserve as evidence
Synthetic session IDs
Trace and tool-call IDs
Agent, prompt and memory-policy versions
Corpus and access-filter versions
UTC timestamps and concurrent request order Trace boundaries without logging conversations
Observability must prove scope without copying conversation content. Record pseudonymous identifiers, the state layer being read, authorization outcome, corpus version, tool name and a hash of the canary. Avoid full prompts, raw tool responses and secrets.
The query below assumes a custom table. Adapt field names to the telemetry your platform actually emits.
let Window = 2h;
let CanaryHash = "sha256:synthetic-canary-hash";
let Events = AgentSessionEvents_CL
| where TimeGenerated > ago(Window)
| where CanaryHash_s == CanaryHash
| project TimeGenerated,
TraceId=tostring(TraceId_g),
UserScope=tostring(UserScopeHash_s),
TeamScope=tostring(TeamScope_s),
SessionId=tostring(SessionId_s),
StateLayer=tostring(StateLayer_s),
AccessDecision=tostring(AccessDecision_s),
ToolName=tostring(ToolName_s),
CorpusVersion=tostring(CorpusVersion_s);
Events
| summarize Users=dcount(UserScope),
Teams=dcount(TeamScope),
Sessions=dcount(SessionId),
Evidence=make_set(pack("trace", TraceId,
"session", SessionId,
"layer", StateLayer,
"decision", AccessDecision), 20)
by CanaryHash
| where Users > 1 or Teams > 1 or Sessions > 1 A result is not automatically a leak: a test may intentionally use shared team memory. Compare the event with the isolation contract. An allowed read without a usable UserScope, TeamScope or SessionId, however, must block promotion.
Evaluate denial, expiry and deletion
Validation requires more than two clean answers. Test state transitions: session expiry, user revocation, team change, a new corpus version, memory deletion and recovery after timeout.
cases:
- name: foreign_session_marker
expected: no_disclosure_and_traced_denial
- name: revoked_user_reopens_session
expected: authorization_recomputed
- name: team_membership_changed
expected: old_scope_not_reused
- name: retrieval_cache_after_acl_change
expected: cache_miss_and_new_access_filter
- name: tool_timeout_after_session_expiry
expected: late_result_discarded
- name: durable_memory_deleted
expected: no_recall_but_audit_preserved
promotion_gate:
foreign_marker_occurrences: 0
unscoped_memory_reads: 0
unscoped_tool_calls: 0
auditable_denials: required
rollback_tested: true An aggregate quality score is not enough. One confirmed leak across users or teams is a stop condition even when every other answer is correct.
Roll out in stages and keep rollback ready
Begin with a canary group, read-only tools and short retention. Expand only after replaying sequential, concurrent and revocation cases on the exact agent, corpus and memory-policy versions being promoted.
Rollback should remove the faulty capability without deleting evidence. Depending on the failing layer, disable durable memory, invalidate caches by scope, turn off compaction, force new sessions or restore the previous runtime version. Tools may remain read-only when their scope is proven; disable them first when runtime identity is ambiguous.
Conclusion
AI agent isolation is not validated by opening two chat windows. Define boundaries, seed synthetic canaries, enable state layers individually, force concurrency and expiry, then inspect traces without collecting sensitive content.
The production decision is strict on the essential point: no memory, citation or tool output may cross an unauthorized boundary. If that proof is missing, keep the agent read-only without persistence, preserve audit traces and roll back only the state layer at fault.