AgentOps: calibrate the evaluator before trusting an agent quality gate
A production runbook for separating agent regression from evaluator drift by freezing outputs, replaying an adjudicated anchor set, measuring disagreement and keeping promotion rollbackable.
Read article