首页 > AI前沿 > Also Small Models Can Reasonably Self-Evaluate Their Confidence

Also Small Models Can Reasonably Self-Evaluate Their Confidence

arXiv机器学习 2026-09-30 18:42 9 阅读 查看原文

This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains.

Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals.

Our results reveal that while accuracy predictably declines with smaller models and more specialized domains, the reliability of self-evaluated confidence remains largely stable across both dimensions.

This independence means the most capable model is not necessarily the best at self-assessing prediction reliability.

These findings suggest that smaller models can achieve reasonable self-assessed confidence despite lower accuracy, making them viable for resource-constrained deployments.