首页 > AI前沿 > Evaluating Test-Time Scaling of General LLM Agents

Evaluating Test-Time Scaling of General LLM Agents

arXiv自然语言 2026-02-22 09:08 5 阅读 查看原文

LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood.

Investigation of LLM Agent Scaling Axes

In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling.

Realistic Benchmark Introduction

We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments.

Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting.

Scaling Test-Time Compute

Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound.

We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling.

Code is publicly available at https://github.com/cxcscmu/General-AgentBench.