AI
AgentOps: diagnose context-window saturation before raising token limits
A production runbook for reconstructing an AI agent's context, isolating history, retrieval and tool outputs, then validating compaction or rolling back.
An internal operations agent works correctly in a short conversation. After several diagnostic steps, it forgets an early constraint, asks for evidence it already collected, or selects a tool with a broader scope than intended. At the same time, input tokens and latency rise. Raising the context limit or switching models can postpone the symptom without fixing repeated history, overbroad retrieval, or full tool outputs injected on every turn.
This runbook starts from a concrete incident: a Microsoft Foundry agent assists a production diagnosis with documentary sources and read or write tools gated by approval. The objective is to reconstruct the context sent to each model call, preserve constraints that must survive compaction, and decide whether to compact history, reduce retrieval, fix a tool contract, or roll back the agent bundle.
Freeze one degraded journey and its baseline
A long conversation alone does not prove saturation. Select a journey that previously succeeded and a reproducible failure. Compare the same objective under the same version bundle with a short context and with a context close to the degradation point.
incident:
started_at_utc: 2026-08-28T13:20:00Z
journey: qualify_application_gateway_502
symptom:
- approval constraint no longer applied
- tool result requested twice
- input tokens and latency increase by turn
release_bundle:
agent: ops-assistant-2026.08.28.2
prompt: prompt-42
model_deployment: <deployment-name>
tool_catalog: tools-18
retrieval_index: runbooks-2026.08.27
context_policy: context-policy-7
comparison:
healthy_trace_id: <trace-id>
degraded_trace_id: <trace-id>
same_user_goal: true
write_tools_disabled_for_replay: true
stop_conditions:
- a write is attempted without the expected approval
- the final answer cannot cite the evidence used
- context construction cannot be tied to a version
- compaction removes an unresolved production action Preserve trace, conversation, response, and tool-call identifiers. Do not automatically paste full content into an incident ticket: prompts, retrieved documents, and tool results can contain sensitive data. Sizes, roles, versions, fingerprints, and identifiers are often enough to locate accumulation.
Reconstruct the context actually sent
The effective window is larger than the history visible in the user interface. A call can include the system prompt, application instructions, a summary, prior turns, retrieved documents, tool definitions, and tool results. The runtime can also add elements the client does not display.
Microsoft Foundry traces can expose inputs, outputs, tool calls, latency, and token consumption. Use that evidence to produce one manifest per model call. When the runtime does not provide an exact token breakdown, measure bytes or characters per segment and estimate offline; never present an estimate as billed usage.
{
"traceId": "<trace-id>",
"responseId": "<response-id>",
"conversationId": "<conversation-id>",
"modelCallId": "<model-call-id>",
"turn": 12,
"versions": {
"agent": "ops-assistant-2026.08.28.2",
"prompt": "prompt-42",
"toolCatalog": "tools-18",
"retrievalIndex": "runbooks-2026.08.27",
"contextPolicy": "context-policy-7"
},
"segments": [
{"type": "system", "items": 2, "bytes": 14800, "digest": "<sha256>"},
{"type": "summary", "items": 1, "bytes": 6200, "digest": "<sha256>"},
{"type": "history", "items": 18, "bytes": 78200, "digest": "<sha256>"},
{"type": "retrieval", "items": 9, "bytes": 112000, "digest": "<sha256>"},
{"type": "tool_output", "items": 6, "bytes": 164000, "digest": "<sha256>"},
{"type": "tool_schema", "items": 14, "bytes": 44800, "digest": "<sha256>"}
],
"usage": {"inputTokens": 0, "outputTokens": 0},
"result": "tool_repeated"
} Compare the first healthy call, the first degraded call, and the call immediately before it. Look for a segment that grows without adding a new decision. An identical tool output carried across turns is a context-construction defect. Different but highly redundant documents point to retrieval. History that retains every raw detail without a synthesized state points to conversation memory.
Measure growth by segment
Normalize manifests into a dedicated table or custom Application Insights events. The following query assumes a custom AgentContextEvents table; map it to the schema your instrumentation actually emits.
let StartTime = datetime(2026-08-28T13:00:00Z);
let EndTime = datetime(2026-08-28T14:00:00Z);
AgentContextEvents
| where TimeGenerated between (StartTime .. EndTime)
| where Journey == "qualify_application_gateway_502"
| summarize
InputTokens=max(InputTokens),
SystemBytes=sumif(SegmentBytes, SegmentType == "system"),
SummaryBytes=sumif(SegmentBytes, SegmentType == "summary"),
HistoryBytes=sumif(SegmentBytes, SegmentType == "history"),
RetrievalBytes=sumif(SegmentBytes, SegmentType == "retrieval"),
ToolOutputBytes=sumif(SegmentBytes, SegmentType == "tool_output"),
ToolSchemaBytes=sumif(SegmentBytes, SegmentType == "tool_schema"),
DistinctDigests=dcount(SegmentDigest),
DurationMs=max(DurationMs),
Results=make_set(Result, 10)
by ConversationId, ModelCallId, Turn, AgentVersion, ContextPolicyVersion
| extend TotalBytes = SystemBytes + SummaryBytes + HistoryBytes
+ RetrievalBytes + ToolOutputBytes + ToolSchemaBytes
| order by ConversationId asc, Turn asc Tokens, latency, and one dominant segment rising together support the diagnosis, but do not prove the functional cause. Replay the same case with that segment bounded. If the agent still fails with a short context, investigate the prompt, model, permissions, or tool itself.
Protect invariants before compacting
Aggressive compaction can make an agent cheaper while making it unsafe. Classify information before reducing it.
Keep these elements in a stable, versioned form:
- current objective, scope, and stop conditions;
- safety constraints and actions that require approval;
- identity, environment, and resource actually targeted;
- completed actions with outcome and idempotency identifier;
- open decisions, hypotheses, and the evidence supporting them;
- references to sources with version and retrieval time.
Compact or replace with a reference: resolved detail, duplicated source material, raw tool outputs after useful fields have been extracted, interface messages, and verbose traces. Do not let a generated summary become the only evidence for a production action; retain an immutable pointer to the source result.
Bound history, retrieval, and tools separately
One global limit makes behavior hard to explain. Apply a budget to each segment and define what happens when it is reached.
version: context-policy-8
budgets:
recent_turns: <derived-from-evaluation>
retrieved_documents: <derived-from-evaluation>
bytes_per_tool_output: <derived-from-tool-contract>
total_input_tokens: <below-model-and-runtime-limit>
history:
keep_recent_turns_verbatim: true
compact_into_state_ledger: true
preserve_open_decisions: true
preserve_approval_and_action_ids: true
retrieval:
deduplicate_by_source_and_section: true
require_source_version: true
reject_unbounded_results: true
tools:
prefer_structured_summary: true
retain_full_result_by_reference: true
paginate_large_collections: true
never_repeat_unchanged_payload: true
on_budget_exhausted:
disable_write_tools: true
return_bounded_diagnostic: true
emit_trace_and_policy_version: true Enforce the budget in the runtime, retrieval component, or tool gateway, not only in the prompt. The model cannot guarantee that an oversized payload will never be sent to it.
Evaluate both process and outcome
Lower token usage is not a sufficient promotion gate. Build a small dataset from redacted traces: short conversation, long history, redundant documents, oversized tool output, stale conflicting instruction, and an action that requires approval.
Measure at least task completion, task adherence, tool selection and success, tool-output utilization, groundedness, latency, input tokens, and compaction count. Foundry evaluators can cover task adherence, tool behavior, and response quality; deterministic controls must separately prove that no write bypasses approval.
cases:
- id: short_baseline
expected: {task_completed: true, context_compactions: 0}
- id: long_history_with_open_decision
expected: {open_decision_preserved: true, repeated_tool_call: false}
- id: redundant_retrieval
expected: {sources_deduplicated: true, grounded_response: true}
- id: oversized_tool_output
expected: {bounded_summary_used: true, full_result_reference_kept: true}
- id: approval_required_after_compaction
expected: {write_before_approval: false, approval_context_preserved: true}
promotion_gates:
- no regression on task completion and task adherence
- no unauthorized write attempt
- no repeated side effect
- bounded input distribution on long cases
- latency remains inside the accepted journey SLO Run nondeterministic evaluations more than once and retain the evaluator model version. Compare the candidate policy with the previous one on the same dataset; an isolated score without a baseline does not support a production decision.
Canary the policy and prepare rollback
Version the prompt, model, tool catalog, retrieval index, summarization strategy, and budgets as one bundle. Enable the candidate policy on read-only journeys first, then on a small share of conversations. Watch tokens per intent, repeated calls, budget denials, task completion, latency, and approval requests.
Keep the policy when the degradation point disappears without losing invariants or business outcomes. Restore the previous bundle when open decisions disappear, tool calls become less reliable, or compaction increases ungrounded answers. Keep write tools disabled when traces cannot prove the state of actions already started.
Conclusion
Context-window saturation is a state-construction problem, not only a model-capacity problem. Reconstruct what is actually sent, attribute growth to history, retrieval, schemas, or tool outputs, and protect invariants before reducing anything.
The production decision then becomes explicit: promote a bounded policy after evaluation and canary, fix the component injecting excess data, or roll back the complete bundle. Raising the limit is defensible only when the need is proven and the growth remains controlled.