首页 > AI前沿 > When KL Regularization Misfires in Group Policy Optimization

When KL Regularization Misfires in Group Policy Optimization

arXiv自然语言 2026-10-08 23:39 4 阅读 查看原文

研究背景

Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates.

分析潜在失败模式

We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards.

提出Zero-Sum Calibrated Policy Optimization (ZCPO)

We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate.

实验结果

Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.