首页 > AI前沿 > On Calibration of Large Language Models: From Response To Capability

On Calibration of Large Language Models: From Response To Capability

arXiv自然语言 2026-02-14 09:07 5 阅读 查看原文

Accurate confidence estimation is critical for reliable use of large language models (LLMs).

Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output.

However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall.

We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability.

To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries.

We formally distinguish capability calibration (CC) from response calibration (RC) and show that the two differ both theoretically and empirically.

We further show that CC is better suited than RC to applications like pass@k prediction and inference budget allocation.

Finally, we evaluate common confidence estimation methods to understand the practical feasibility of CC.