首页 > AI前沿 > Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

arXiv自然语言 2026-09-29 17:28 6 阅读 查看原文

Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead.

Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes.

We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning.

Findings

First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence.

Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes.

Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels.

Introducing RFPO

Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes.

Results

We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias.

Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead.

This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost:

because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete.

Conclusion

Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.