首页 > AI前沿 > Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning

Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning

arXiv自然语言 2026-10-01 12:00 8 阅读 查看原文

Global structured pruning requires channels from different layers to compete under a shared sparsity budget, raising two coupled challenges:

  • identifying which channels should be retained
  • making their scores comparable across layers

Raw channel scores can contain block-common scale that leaves within-block ordering unchanged but distorts model-wide competition.

Our experiment indicates that similar layer-wise allocations can retain substantially different FFN channels, so layer allocation alone does not determine channel identity.

Motivated by this separation, we introduce Global Relative Kinetic Utility (Global RKU), a label-free criterion that separates channel importance estimation from cross-layer comparison.

Global RKU measures channel participation using a final-hidden-state activation-gradient signal, then applies block-relative normalization to mitigate block-common scale while preserving within-block ordering, requires only unlabeled calibration inputs, and produces a static pruning topology in a single calibration stage.

Under questions-only calibration on Qwen-2.5-7B, RKU-GISP Mean3 margins are -0.98, +3.79, and +8.61 points at 30%, 40%, and 50% sparsity, respectively (average +3.81).

Additional Qwen evaluations cover non-mathematical reasoning, recovery, held-out transfer, and physical deployment.

Separately, replacing Wiki16K with questions-only Q16K improves RKU's Mean3 at every tested sparsity on Qwen, Llama, and Gemma.

Our ablation study shows relative-normalization gains of 14.42 and 5.53 Mean3 points at 40% and 50% sparsity, respectively; the common-seed audit is positive in all 27 seed-task comparisons.