首页 > AI前沿 > ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

arXiv自然语言 2026-10-08 14:28 4 阅读 查看原文

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.

Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness.

We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning.

RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions.

Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons.

Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery.

These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.