首页 > AI前沿 > Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

arXiv自然语言 2026-10-03 11:01 3 阅读 查看原文

Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations.

Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as conformal prediction.

To address this issue, we propose a paraphrase-aware UQ framework robust to semantic rewordings.

Our approach trains a lightweight proxy model on LLM hidden states and aggregates its predictions across paraphrases to construct label-wise nonconformity scores.

Under score exchangeability, conformal calibration retains marginal coverage.

This guarantee can also hold under test-only rewording, provided that the paraphrase pipeline satisfies an additional distributional alignment condition.

Evaluation Settings

We evaluate three settings (normal, fully reworded, and semi-reworded) which apply rewording to neither dataset, both calibration and test datasets, or only the test dataset, respectively.

Results

Across seven multiple-choice QA benchmarks and multiple model families, our method produces compact prediction sets with empirical coverage generally near the nominal target, even in the semi-reworded setting.

Ablation Studies

Ablation studies show that the learned proxy accounts for most of the reduction in set size, while paraphrase-augmented training and inference-time aggregation improve stability under rewording.

Code is available at https://github.com/Raina-Xin/PA_Score.