首页 > AI前沿 > Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

arXiv自然语言 2026-09-28 04:16 7 阅读 查看原文

Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers.

We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run.

Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop.

A calibration test then selects a schedule at a stated selective-risk target.

In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits.

Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier.

At matched coverage, it keeps the fewest wrong answers among these scorers.

Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage.

These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers.

Code is available at https://github.com/js-lee-AI/PriceCheck.