Can an AI agent evaluation pass if a human quietly rescued the run?
Can an AI agent evaluation pass if a human quietly rescued the run?
A final answer can be correct while the trajectory is expensive, unsafe or impossible to reproduce. I think agent evaluations need to separate at least four dimensions: outcome, trajectory, human intervention and recovery.
The case itself should be frozen: source snapshot, task packet, allowed tools, permission boundaries, expected effects and acceptance checks. The run receipt should record tool calls, retries, cost, approvals and every durable side effect. Then the same case can be replayed after a model, prompt, tool or policy change.
Disclosure: I’m building ChaseOS Studio, where this kind of evidence and approval boundary is part of the control-plane design. I wrote up the practical evaluation loop here: https://chaseos.ai/blog/reproducible-model-evaluations-ai-agents?utm_source=producthunt&utm_medium=forum&utm_campaign=agent_governance_series
What do you treat as a failed run even when the final output looks correct?

Replies