首页 > AI前沿 > When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

arXiv自然语言 2026-10-08 13:54 4 阅读 查看原文

On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models.

In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?

We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency.

Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup.

We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios.

Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher.

For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal.

In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens.

Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.