首页 > AI前沿 > Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

arXiv自然语言 2026-09-29 13:07 6 阅读 查看原文

Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle.

In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors.

We propose CaRE-KD

CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization.

CaRE-KD has two components:

  • a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence,
  • a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student.

We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce.

Empirically

Across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, $α$--$β$ divergence).

Highlights include:

  • up to $+3.2$ average ROUGE-L on instruction-following tasks,
  • $+2.1$ pass@1 on MBPP,
  • $+1.7$ accuracy on GSM8k,
  • $+1.8$ accuracy on CollegeMath over the strongest baseline,
  • with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL).

Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.