首页 > AI前沿 > CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

arXiv自然语言 2026-10-06 01:57 6 阅读 查看原文

Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment.

We introduce CLIFT, a training and test-time scaling method built around conformal self-verification.

During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline.

At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge.

This single mechanism supports three settings.

WebArena Infinity

On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents.

VisualWebArena

On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness.

Online Mind2Web

On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation.

Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.