首页 > AI前沿 > Learning from Think-Mode Advantage via On-Policy Distillation

Learning from Think-Mode Advantage via On-Policy Distillation

arXiv自然语言 2026-09-29 16:51 6 阅读 查看原文

Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode.

We study learning from this think-mode advantage via on-policy distillation (OPD).

OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes.

Privileged reasoning is used during distillation rather than student inference.

Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response.

Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when responses reach the same outcome.

We summarize this interaction with trace-response divergence (TRD) and introduce ThinkOPD, which routes supervision at the response level by combining group-relative reward gain with a TRD-based compatibility proxy.

Final response weights are normalized within each rollout group.

Across mathematical reasoning and code generation, ThinkOPD outperforms Uniform ThinkOPD in both same-model settings and both cross-model teacher-student pairs, and it exceeds representative rationale and self-distillation baselines in a controlled comparison.

Controlled interventions show that outcome benefit and the TRD-based proxy provide complementary routing signals in this setting.

Think-enabled OPD provides a controlled setting for studying how teacher advantage becomes transferable along student responses.