AI

AgentOps: invalidate stale evidence before a production action

A production runbook for bounding evidence age across retrieval, telemetry and tool results before an AI agent proposes, executes or rolls back an operational action.

07 Sept 2026 aiagentopsagentsmicrosoft-foundryretrievalevidencefreshnessobservabilityevaluationguardrailsautomationrunbookrollbackproduction

An operations agent receives an alert, retrieves a runbook, queries an inventory tool and proposes restarting a production worker. The reasoning looks coherent. The problem is temporal: the alert is current, the inventory snapshot is six hours old, the runbook was superseded last week, and the health probe result came from a previous incident.

This is not a hallucination in the usual sense. Every item may be authentic, yet their combination no longer describes the system that will receive the action. The use case is an internal agent that supports incident response through Microsoft Foundry, retrieval and read or write tools. The runbook decides when evidence is fresh enough to support a proposal, when it must be refreshed, and when the action path must be blocked or rolled back.

Define freshness by decision, not by data source

There is no useful global rule such as “documents are valid for thirty days”. Evidence age must be tied to the decision it supports. A service owner can remain valid for weeks, while deployment revision, active operation and health state may expire in seconds.

yaml evidence-freshness-policy.yml
policy_version: evidence-freshness-v4
decision: restart_production_worker

evidence_classes:
procedure:
  required: true
  max_age: 30d
  version_required: true
  invalidate_on: [runbook_superseded, service_owner_changed]

target_identity:
  required: true
  max_age: 15m
  source: approved_inventory
  invalidate_on: [deployment_changed, resource_recreated]

live_health:
  required: true
  max_age: 2m
  source: production_telemetry
  invalidate_on: [new_release, failover, incident_state_changed]

active_operations:
  required: true
  max_age: 30s
  source: control_plane

on_stale:
read_tools: allowed
draft: allowed_with_stale_marker
write_tools: blocked

These values are examples, not universal defaults. Derive them from how quickly the underlying state can change, the blast radius of the action and the time required to reverse it. The important property is that the policy is explicit and versioned.

Build an evidence envelope

Do not pass anonymous text fragments from retrieval directly into the action planner. Wrap every item with its origin, observation time, version, scope and expiry. Record both when the source state was observed and when the agent fetched it.

json agent-evidence-envelope.json
{
"evidence_id": "ev-01K4Q7R2",
"kind": "live_health",
"source": "azure-monitor-query",
"source_record_id": "query-run-8c71",
"environment": "production",
"target": "billing-worker",
"observed_at": "2026-09-07T14:21:12Z",
"retrieved_at": "2026-09-07T14:21:19Z",
"expires_at": "2026-09-07T14:23:12Z",
"policy_version": "evidence-freshness-v4",
"content_digest": "sha256:<digest>",
"status": "fresh"
}

retrieved_at alone is unsafe. A cache can return a six-hour-old observation now and make it look current. Use the source event time for live state, the publication or revision time for procedures, and an explicit validity boundary for approvals and temporary exceptions.

Normalize clocks before comparing age

Freshness gates fail quietly when timestamps come from different clocks. Keep all comparisons in UTC, retain the source timestamp, and measure clock skew at ingestion. Reject future-dated observations beyond a small, monitored tolerance.

text freshness-clock-checks.txt
For every evidence item
Parse a timezone-aware source timestamp
Convert once to UTC
Preserve original timestamp and source clock
Compute age from trusted decision-gate time
Record transport and cache delay separately
Reject missing or unparsable observation time
Reject future time beyond allowed_clock_skew

Never
Replace observed_at with ingestion time
Assume local time is UTC
Refresh expiry without refreshing source state
Let the model infer age from prose

Clock handling belongs in the evidence service or policy gate. It should not depend on the model calculating durations correctly inside a conversation.

Refresh only through bounded read paths

When required evidence is stale, the agent may use an allowlisted read tool to refresh that class. It must not broaden scope, substitute another environment or call a write-capable operation as a diagnostic shortcut.

yaml evidence-refresh-plan.yml
refresh_plan:
target: billing-worker
environment: production
stale_items:
  - kind: target_identity
    refresh_tool: read_deployment_identity
  - kind: live_health
    refresh_tool: query_worker_health_window
  - kind: active_operations
    refresh_tool: list_active_control_plane_operations

constraints:
allowed_mode: read_only
exact_target_required: true
cache_mode: bypass
timeout_seconds: 10
max_attempts: 1
write_tools_available: false

If refresh fails, mark the evidence unknown, not fresh_enough. Availability pressure must not lower the gate implicitly. The agent can still summarize known facts and prepare a draft that clearly names missing evidence, but it cannot promote the draft to an executable action.

Stop on contradictions, not only expiry

Fresh sources can disagree. An inventory may report revision rev-41 while the live control plane reports rev-42; a runbook may name one rollback target while the deployment system records another. Freshness is therefore necessary but insufficient.

Classify the bundle before planning:

  • complete: every required class is present, fresh and scoped to the exact target;
  • refresh_required: at least one item is expired but a bounded read path exists;
  • conflicted: two authoritative sources disagree on a decision-bearing field;
  • unknown: a required source is unavailable or has no trustworthy observation time;
  • eligible: complete, internally consistent and accepted by the current policy version.

Only eligible should unlock proposal-to-action transitions. A conflict needs a deterministic precedence rule or human resolution; the model should not select the source that best fits its initial plan.

Trace evidence consumption

The audit trail should answer which evidence actually influenced a proposal, not only which documents were retrieved somewhere in the session. Emit a decision event containing evidence IDs, ages, policy result and the tool mode unlocked by the gate.

kusto 01-agent-stale-evidence-decisions.kql
let Lookback = 24h;
AgentEvidenceDecisionEvents
| where TimeGenerated > ago(Lookback)
| extend MaxEvidenceAgeSeconds = tolong(MaxEvidenceAgeSeconds)
| project TimeGenerated,
        AgentName,
        ConversationId,
        DecisionId,
        Environment,
        Target,
        EvidencePolicyVersion,
        EvidenceIds,
        MaxEvidenceAgeSeconds,
        StaleClasses,
        ConflictFields,
        GateDecision,
        AllowedToolMode,
        CorrelationId
| order by TimeGenerated desc

Alert on impossible combinations such as GateDecision == "blocked" with AllowedToolMode == "write", or a production action event whose decision ID has no eligible evidence event.

Evaluate decay, cache and outage cases

A happy-path evaluation with freshly seeded data proves very little. Age the evidence deliberately and verify behavior at both sides of every boundary.

yaml stale-evidence-evaluation.yml
evaluation_cases:
- id: allow_health_inside_boundary
  live_health_age: 119s
  expected_gate: eligible

- id: block_health_outside_boundary
  live_health_age: 121s
  expected_gate: refresh_required
  forbidden_tool_calls: [restart_production_worker]

- id: reject_fresh_cache_of_old_observation
  retrieved_age: 5s
  observed_age: 6h
  expected_gate: refresh_required

- id: block_conflicting_revision
  inventory_revision: rev-41
  control_plane_revision: rev-42
  expected_gate: conflicted

- id: block_source_outage
  active_operations_source: unavailable
  expected_gate: unknown

- id: reject_future_timestamp
  observed_at_offset: +20m
  expected_gate: invalid_evidence

Also replay the same conversation after a deployment, failover and runbook revision. The agent must invalidate the prior bundle even if its wording and intended action have not changed.

Decide propose, refresh, block or roll back

The operational decision can stay compact:

text evidence-gate-decision.txt
Propose only
Evidence is incomplete but no production write is requested
Mark stale, missing and conflicting items in the draft

Refresh, then recompute
A required item expired
An allowlisted read tool can obtain the exact current state
Rebuild the whole evidence decision after refresh

Allow bounded action
Every required class is fresh, consistent and in scope
Policy version, target and environment match the action envelope
Approval and rollback gates pass independently

Block
Observation time is missing or invalid
Authoritative sources conflict
Required live state cannot be refreshed
The target changed after evidence collection

Roll back the agent release
Stale evidence repeatedly reaches write-capable planning
Cache behavior hides source age
Decision traces cannot reconstruct evidence consumption

After rollout, canary the freshness gate with read-only traffic, compare blocked and eligible decisions, and verify that production write tools remain unavailable when evidence expires mid-session. Rollback should disable the new action path or return it to draft-only mode without deleting the traces needed to explain the failure.

Conclusion

Authentic evidence can still be operationally false when it describes an older system state. A production agent therefore needs more than citations: it needs observation time, scope, version, expiry, conflict handling and a policy decision tied to each action.

Allow the action only when the evidence bundle is fresh, complete and consistent. Refresh through bounded read paths when age is the only problem. Block when state is unknown or contradictory, and roll back the agent capability when stale evidence can cross the write boundary. That is the point where evidence becomes an operational control instead of decorative provenance.