首页 > AI前沿 > Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

arXiv机器学习 2026-10-06 04:31 6 阅读 查看原文

We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion.

Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored).

Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run.

Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown.

We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation that satisfies hard resource constraints and adaptively reallocates unused budget.

We show how to use HARP to construct lower predictive bounds (LPBs) on the time-to-event and to estimate evaluation metrics such as the jailbreak rate on a fixed benchmark.

Although HARP induces dependence in acquisition decisions across different trajectories, we prove that HARP never exceeds the target budget, that its LPBs have finite-sample coverage guarantees, and that its metric estimates are unbiased.

Experiments on agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations show that HARP achieves coverage close to the nominal level with low variance, while never exceeding the given budget.