首页 > AI前沿 > Efficient Best-of-N policy evaluation for inference-time alignment

Efficient Best-of-N policy evaluation for inference-time alignment

arXiv机器学习 2026-10-07 08:26 5 阅读 查看原文

Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.

Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods.

In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods.

We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone.

We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets.

We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator.

Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules:

  • maximizing the estimated policy value
  • maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee

Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.