首页 > AI前沿 > From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation

From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation

arXiv自然语言 2026-09-29 21:06 5 阅读 查看原文

On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher.

Teacher interventions can improve these trajectories, but also change the distribution on which the student learns.

Our controlled studies show

that rollout quality alone is an incomplete criterion for allocating teacher guidance.

Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load.

In a training probe with a restricted rollout horizon

peak student accuracy and performance retention favor different intervention strengths.

The preferred intervention depth and placement also vary across benchmarks.

These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates.

Its {policy disagreement score} combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs.

Across eight mathematical reasoning benchmarks

MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark.

MAESTRO also reduces average training response length by 67.3% relative to standard OPD.

The code is available at https://github.com/yhao-wang/MAESTRO.