首页 > AI前沿 > COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

arXiv机器学习 2026-10-07 15:43 5 阅读 查看原文

Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.

Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective.

We show that this policy-side correction alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term advantage staleness.

We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance.

This motivates the hypothesis that policy- and advantage-side correction should be coordinated.

We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation.

Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other.

COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting.

It also offers a broad high-performing parameter region and improved training stability.

In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training.

These gains persist at 64-step policy staleness.

COPC adds minimal step-time overhead over asynchronous PPO and retains a $1.7\times$ step-time speedup over synchronous PPO.