首页 > AI前沿 > Uncertainty-Normalized Margins for Direct Preference Optimization

Uncertainty-Normalized Margins for Direct Preference Optimization

arXiv机器学习 2026-09-30 07:08 6 阅读 查看原文

Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback.

We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale.

Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy.

Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale.

For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable.

We introduce a practical procedure for learning the scale.

Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length.

We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge.

On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO.

These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.