首页 > AI前沿 > TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

arXiv自然语言 2026-10-08 18:52 4 阅读 查看原文

Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability.

It may instead reflect a change in the evaluation.

We introduce TRACE

We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible.

Controlled suite of 25 synthetic tasks

In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent.

Public $τ^2$-bench tasks

On public $τ^2$-bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate.

Follow-up study

In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within $\pm$0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects.

Identical reruns

Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation.

Frontier LLM judges

Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome.

TRACE

TRACE thus separates what a score change says about the agent from what it says about the measurement.