Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own.
The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy.
We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation.
The ratio is therefore restored exactly, from log-probabilities PPO already computes.
Autoregressive generation and canonical-order constructive optimization are both history-injective.
The exact correction pays importance-sampling variance that grows with the horizon, so we generalize it to a one-parameter family with PPO ($α{=}0$) and the full correction ($α{=}1$) as endpoints: a single bias--variance knob.
A gradient-level analysis of the unclipped surrogate identifies two channels the correction acts through and three conditions under which it carries signal; an enumerable testbed confirms the conditions' predictions.
On hard credit-assignment scheduling tasks, a short corrected warmup with aggressive early sample reuse learns faster than PPO and than the same reuse uncorrected; the marginal gain grows with task difficulty ($+0.02$ to $+0.09$ learning-curve AUC), and the early win over PPO tracks the prefix bias that reuse incurs.
A correction held throughout, or applied where clipping already contains the reuse bias, is null to harmful.