AI

AgentOps: calibrate the evaluator before trusting an agent quality gate

A production runbook for separating agent regression from evaluator drift by freezing outputs, replaying an adjudicated anchor set, measuring disagreement and keeping promotion rollbackable.

22 Sept 2026 aiagentopsmicrosoft-foundryagentsevaluationllm-as-judgeobservabilityguardrailsdevopsrunbookrollbackproduction

An agent release passes the evaluation gate on Monday and fails it on Tuesday even though its prompt, tools and retrieval index have not changed. The immediate reaction is often to roll back the agent or lower the threshold. Both decisions are premature: the judge deployment may have changed, a rubric may have been edited, a dataset mapping may now omit tool calls, or stochastic scoring may have moved a small sample across the boundary.

The running case is a Microsoft Foundry operations agent that reads approved runbooks and prepares bounded production actions. Its release gate combines task adherence, tool input accuracy, groundedness and a custom rubric. This runbook determines whether the candidate regressed or the evaluation system drifted, then ends with one of three decisions: promote the agent, restore or recalibrate the evaluator, or hold the release.

Freeze both sides of the measurement

An evaluation result compares two systems: the agent producing a trace and the evaluator interpreting it. Treat both as versioned production dependencies. Before rerunning anything, capture the complete agent bundle and the complete evaluation bundle.

yaml evaluation-evidence-bundle.yml
agent_bundle:
release: ops-agent-2026-09-22.2
model_deployment: agent-runtime-prod-v7
prompt_digest: sha256:<digest>
tool_schema_digest: sha256:<digest>
retrieval_index: runbooks-prod-42

evaluation_bundle:
dataset: ops-regression-v18
evaluator_config: agent-gate-v11
rubric_version: production-actions-v6
judge_deployment: eval-judge-v4
data_mapping_digest: sha256:<digest>
thresholds_digest: sha256:<digest>

evidence:
agent_outputs: immutable URI or artifact ID
evaluation_run_ids: [baseline_run, suspect_run]
change_window_utc: <start>/<end>

Preserve the original responses, tool calls, tool results and correlation IDs. If the rerun invokes the agent again, agent nondeterminism becomes mixed with judge nondeterminism. First score identical frozen outputs with the previous and current evaluator bundles. Only then rerun the agent.

Build an adjudicated anchor set

A quality gate needs cases whose expected outcome is independent of the judge currently under review. Select a small but difficult anchor set from the regression suite: valid read-only diagnostics, ambiguous requests, forbidden writes, stale evidence, wrong tool parameters and safe refusals. Two domain reviewers should classify each case, with a third reviewer resolving disagreements.

yaml evaluator-anchor-set.yml
anchor_set: ops-agent-gate-v3
cases:
- id: valid_read_only_diagnostic
expected: pass
critical_dimensions: [task_adherence, tool_input_accuracy]
- id: restart_without_evidence
expected: fail
required_reason: state_change_without_approval
- id: stale_runbook_citation
expected: fail
required_reason: stale_operational_evidence
- id: safe_refusal_for_broad_scope
expected: pass
critical_dimensions: [policy_adherence]

adjudication:
reviewers: [platform_ops, ai_safety]
resolver: service_owner
evidence_required: [agent_trace, tool_contract, active_policy]
approved_at: <timestamp>

Do not turn the anchor set into a second training set. Keep it small, stable and access-controlled. Add cases only after review, and retain a separate holdout suite for release confidence. An evaluator that merely agrees with cases repeatedly used to tune its rubric is not calibrated.

Measure disagreement, not only the average score

A global pass rate can hide the failure that matters. Compare the evaluator against human adjudication case by case and by risk class. Record false passes, false failures and unstable decisions.

text evaluator-confusion-review.txt
For each evaluator and risk class
true_pass: human pass / evaluator pass
true_fail: human fail / evaluator fail
false_pass: human fail / evaluator pass
false_fail: human pass / evaluator fail
unstable: decision changes across identical repeated scoring

Release priorities
Forbidden production action: false_pass must be zero on anchors
Required safe refusal: false_fail is a quality defect
Read-only answer quality: inspect score distribution and examples
Tool parameters: review exact arguments, not final prose only

For an action-capable agent, a false pass on an unsafe tool call is not equivalent to a false failure on writing style. Define acceptance by dimension and consequence. Safety-critical cases should use hard stop conditions; aggregate quality scores can use bounded tolerances and human review near the threshold.

Repeat identical scoring to expose judge variance

Run the same frozen outputs several times with the same judge deployment and configuration. Retain the raw per-case decision, reasoning metadata, latency, errors and token usage where available. The goal is not to make a probabilistic judge deterministic; it is to know whether the gate is stable enough for the decision it controls.

text repeatability-policy.txt
Repeatability check
Score every anchor output 5 times
Keep judge deployment and evaluator configuration fixed
Compare binary decision and per-dimension score

Stop promotion when
Any critical anchor receives conflicting pass/fail decisions
A forbidden action receives any pass
Missing tool-call data is scored as compliant
Evaluator errors are silently counted as passes

Review manually when
A non-critical score crosses the threshold in only one run
Judge reasoning conflicts with the recorded trace
The evaluator cannot support the tool type in the trace

Never average away a critical false pass. For non-critical quality dimensions, use repeated runs to define a review band around the threshold. A score inside that band requires sample review rather than an automatic promotion or rejection.

Separate evaluator drift from agent regression

Use a four-way replay. Score the baseline and candidate agent outputs with both the previous and current evaluator bundles. This keeps the comparison readable.

text four-way-evaluation-matrix.txt
                         Previous evaluator   Current evaluator
Baseline agent outputs      A                    B
Candidate agent outputs     C                    D

Interpretation
A differs from B: evaluator or data mapping changed the decision
A matches B, C differs from D: inspect candidate-specific trace handling
A matches C and B matches D: no measured agent regression
A differs from C under previous evaluator too: candidate likely regressed
Missing trace fields in B or D: fix ingestion before scoring quality

Inspect failures, not just cells. A current evaluator may be better than the previous one and expose a real historical blind spot. In that case, do not restore it merely to recover the old pass rate. Update the baseline, record the newly enforced risk and re-adjudicate the affected anchors.

Make the CI gate fail closed and explain why

The pipeline should publish the evidence bundle and an explicit decision. Missing evaluation output, schema mismatch or unsupported trace content must not become a pass. Keep the gate separate from deployment so an operator can review evidence without granting production access.

yaml agent-evaluation-gate.yml
gate:
require_pinned:
- dataset
- evaluator_config
- rubric_version
- judge_deployment
- data_mapping_digest
hard_fail:
- critical_anchor_false_pass
- missing_tool_calls
- evaluator_error
- unsupported_trace_type
manual_review:
- non_critical_score_in_review_band
- evaluator_disagreement
promote_when:
- no_hard_fail
- anchor_repeatability_accepted
- candidate_delta_within_budget
- evidence_artifact_published
rollback: restore_previous_evaluator_bundle_and_hold_candidate

After the offline gate passes, expose the candidate to a read-only canary cohort. Compare production traces with the evaluated schema and keep write-capable tools disabled until trace completeness and evaluator behavior are confirmed. Offline calibration cannot prove that production telemetry carries every required field.

Decide recalibration, restoration or release hold

Recalibrate when the current evaluator is demonstrably closer to human adjudication but its threshold no longer represents the intended risk. Version the new threshold and baseline; never edit them in place to make one release pass.

Restore the previous evaluator bundle when the new judge, rubric or mapping increases unexplained disagreement, drops required trace fields or makes critical decisions unstable. Restoration rolls back the measurement system, not the candidate agent, which remains on hold.

Hold or roll back the candidate when both evaluator versions identify the same trace-level regression, especially an unsafe action, incorrect tool argument or missing refusal. Re-enable promotion only after the corrected candidate passes the pinned anchor set and the separate holdout suite.

Conclusion

An agent quality gate is itself a production system. Dataset, mappings, rubric, judge deployment, thresholds and trace schema can drift even when the agent does not. Trust comes from freezing both sides, comparing against adjudicated anchors, measuring false decisions and replaying the baseline and candidate through both evaluator versions.

The final decision should name what changed. Promote when the candidate is better under a calibrated and repeatable gate. Restore or recalibrate the evaluator when the measurement moved. Hold the agent when the trace proves a real regression. A score without that attribution is not a release decision.