首页 > AI前沿 > Calibration as a First-Class Criterion in LLM Evaluation

Calibration as a First-Class Criterion in LLM Evaluation

arXiv自然语言 2026-09-22 22:28 6 阅读 查看原文

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP.

Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful.

We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation.

Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it.

Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment.

Most benchmarks in use today already provide both, meaning calibration can be reported immediately.

For open-ended generation, however, defining these two inputs is still an open challenge.

We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.