Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.
Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective.
We show that this policy-side correction alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term advantage staleness.
We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance.
This motivates the hypothesis that policy- and advantage-side correction should be coordinated.
We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation.
Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other.
COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting.
It also offers a broad high-performing parameter region and improved training stability.
In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training.
These gains persist at 64-step policy staleness.
COPC adds minimal step-time overhead over asynchronous PPO and retains a $1.7\times$ step-time speedup over synchronous PPO.