首页 > AI前沿 > ReTaCo: Residual-Target Control for On-Policy Distillation

ReTaCo: Residual-Target Control for On-Policy Distillation

arXiv机器学习 2026-09-30 16:27 8 阅读 查看原文

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly.

Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-$k$ probabilities to limit cost.

Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary.

We prove that the resulting loss keeps pushing the student's top-$k$ mass toward one even after the student matches the teacher's relative probabilities within the top-$k$ set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability.

We propose ReTaCo (Residual-Target Control), which keeps the top-$k$ tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL.

At a fixed prefix, we prove that the population objective has a unique optimum whose top-$k$ mass lies between $m$ and $m+β(1-m)$ and increases monotonically with $β$; at $β=0$, underestimated top-$k$ tokens still receive non-vanishing recovery gradients.

Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.