AI
AgentOps: detect evaluation set contamination before promoting an agent
A production runbook for proving that an AgentOps score is not inflated by leakage across evaluations, prompts, retrieval or tuning data before promoting an AI agent.
An agent candidate gains twelve points on the reference evaluation set. Its answers are more precise, tool calls are better scoped and sensitive refusals are more consistent. Promotion looks straightforward. Then a reviewer notices that several cases closely match examples recently added to the system prompt, while some ground-truth documents are also searchable through the pre-production retrieval index.
The question is no longer whether the agent scores well. The team must determine whether it generalizes to new situations or recognizes answers it has already seen. The use case is an operations agent that searches runbooks, triages incidents and prepares actions for approval. Before promotion, the team needs evidence that evaluation cases remain separate from prompts, tuning data, few-shot examples, document indexes and recycled traces.
Freeze the candidate and its dependencies
A contamination investigation only works if the system under review stops moving. Freeze the complete bundle: model, prompt, tools, guardrails, retrieval index, evaluation configuration and dataset versions. Quarantining a candidate does not interrupt production. The active release keeps serving users while the candidate is held back.
candidate_release: ops-agent-2026-08-25.1
promotion_status: quarantined
bundle:
model_deployment: candidate-v7
prompt_digest: sha256:<prompt-digest>
tool_contract: tools-v18
guardrail_policy: guardrails-v11
retrieval_snapshot: runbooks-2026-08-24
evaluator_version: eval-pipeline-v9
datasets:
development: agent-dev-v12
regression: agent-regression-v8
holdout: agent-holdout-v5
red_team: agent-redteam-v4
forbidden_until_decision:
- change_prompt
- rebuild_retrieval_index
- relabel_failed_cases
- publish_candidate
- overwrite_evaluation_traces Retain the raw evidence as well: normalized inputs, outputs, retrieved sources, tool calls, guardrail decisions and evaluator results. Editing the prompt immediately may hide the leak without answering the core question: how did the candidate know this case?
Map every leakage path
The usual train-test split is not enough for an agent. A case may reach the candidate without being part of model training. It can be copied into the prompt, stored in the retrieval corpus, turned into a demonstration from a production trace, embedded in a tool fixture or exposed through feedback from a previous evaluation run.
Surfaces to compare
development and regression cases
promotion holdout
red-team scenarios
system prompt and few-shot examples
retrieval corpus and metadata
simulated tool fixtures
recycled production traces
fine-tuning or preference data
evaluator outputs and human comments
Critical leakage
expected answer is available to the candidate runtime
case identifier appears in prompt, index or fixture
holdout document was indexed before evaluation
red-team scenario became a refusal example
same logical incident exists in development and holdout
Similarity requiring review
shared template with different values
same runbook but different symptom and decision
automatic paraphrase of an existing case
translated duplicate across English and French corpora Treat English and French as one risk space. Translating a development case and placing the English version in the holdout does not create an independent scenario. Resource names, error codes, runbook steps and expected decisions can preserve the link even when the wording changes.
Attach usable lineage to every case
A filename is not enough provenance. Each case needs an origin, ingestion date, scenario family, deduplication group and a record of every surface that has seen it. The group connects translations, paraphrases and variants derived from the same logical incident.
{
"caseId": "eval-inc-0427-en",
"scenarioFamily": "tool-authorization-failure",
"language": "en",
"sourceClass": "synthetic-reviewed",
"sourceRef": "scenario-pack-2026-08",
"createdAt": "2026-08-12T09:00:00Z",
"dedupGroup": "incident-logic-0427",
"split": "holdout",
"expectedDecision": "diagnose-identity-before-permission-change",
"exposure": {
"prompt": false,
"retrieval": false,
"toolFixture": false,
"fineTuning": false,
"previousEvaluationFeedback": false
},
"contentDigest": "sha256:<normalized-case-digest>",
"reviewStatus": "approved"
} The digest catches exact copies but does not replace semantic deduplication. Normalize whitespace, case, variable identifiers and timestamps, then compare scenario groups and expected decisions. High lexical similarity is a review signal, not automatic proof of leakage.
Scan before running the evaluation
Run contamination checks when assembling the set and again against the actual bundle deployed to pre-production. Compare exact digests, long fragments, rare identifiers and semantic similarity. Inspect the rendered prompt and documents that the runtime can retrieve, not only their source repositories.
blocking:
- exact_match_with_prompt_example
- holdout_document_in_retrieval_snapshot
- expected_answer_in_tool_fixture
- same_dedup_group_across_development_and_holdout
- red_team_case_used_as_candidate_demonstration
review_required:
- semantic_similarity_above_review_threshold
- shared_rare_identifiers
- translated_or_paraphrased_candidate
- provenance_missing
- split_changed_after_result_review
allowed_with_evidence:
- shared_public_runbook_with_distinct_incident_state
- common_operational_template_with_independent_values
- repeated_policy_rule_testing_different_decision_boundaries
outputs:
- immutable_scan_report
- excluded_case_manifest
- reviewer_decisions
- clean_holdout_digest Never let a similarity threshold silently delete cases. It should bring candidates together for review, after which a reviewer decides whether they measure memorization, generalization or a genuinely different capability. Record every exclusion. Otherwise, a team can improve the headline score simply by removing difficult scenarios after seeing the result.
Read scores through lineage
Contamination does not always produce a perfect score. It may appear as an improvement concentrated on older cases, scenarios available through retrieval or families that drove recent prompt changes. Segment results by origin, creation date, deduplication group and suspected exposure class.
AgentEvaluationResults
| where TimeGenerated > ago(14d)
| where CandidateReleaseId == "ops-agent-2026-08-25.1"
| summarize
Cases = dcount(CaseId),
PassRate = 100.0 * countif(Outcome == "pass") / count(),
ToolBoundaryFailures = countif(CheckName == "tool_boundary" and Outcome == "fail"),
MedianCreatedAgeDays = percentile(datetime_diff("day", TimeGenerated, CaseCreatedAt), 50)
by Split, SourceClass, ExposureClass, ScenarioFamily
| order by Split asc, PassRate desc AgentEvaluationResults is a normalized example to adapt to the real telemetry schema. Look for discontinuities: excellent performance on exposed cases but ordinary results on a recent holdout, no progress on a new family, or success correlated with a retrieved document that nearly contains the expected decision. Compare the active release too. A biased set can favor both agents and remain invisible in the candidate delta.
Rebuild a defensible holdout
When contamination is confirmed or cannot be ruled out, moving affected files to a new folder is not enough. Build a holdout from independent sources, cases created after the candidate freeze, or scenarios written by reviewers who cannot see detailed candidate failures. For high-risk decisions, separate the people improving the prompt from those validating the new reference set.
Preserve critical families in the replacement set: refusal to act without evidence, ambiguous scope, tool failure, conflicting sources, missing approval and rollback. Vary details that test generalization, including symptom order, language, resource names and irrelevant context. Do not move the expected operational boundary merely to make the wording different.
Replay both the active release and candidate against the clean holdout. A genuinely stronger candidate should retain a meaningful share of its advantage, especially on blocking invariants. A sharp collapse indicates that the original score mostly measured corpus exposure.
Decide promotion, quarantine or rollback
Promotion can resume when lineage is complete, blocking scans are clean, an independent holdout confirms the gain and sensitive cases have been reviewed. The decision record should identify the exact bundle, the clean dataset digest and every accepted exclusion.
Keep the candidate quarantined when leakage is likely but not localized, retrieval documents are not versioned or the new holdout lacks coverage. Roll back the evaluation pipeline when contaminated results have already selected a model, changed a prompt or justified broader tools. Return to the last release promoted on a defensible dataset, then rebuild the evaluation chain before attempting another promotion.
Rollback is broader than restoring the active model. Remove holdout cases from indexes and examples, invalidate contaminated reports while retaining their audit trail, and stop promotion automation from reusing them.
Conclusion
An AgentOps score is evidence only when the candidate could not study the exam. For an agent connected to sources and tools, that boundary crosses prompts, retrieval, fixtures, traces, tuning data and evaluation feedback.
The final decision must remain operational: promote on an independent, traceable holdout; maintain quarantine while provenance is incomplete; or roll back the selection pipeline when contamination has already influenced a release. The goal is not to protect a benchmark. It is to stop an apparently improved agent from discovering its limits during the first truly unfamiliar incident.