Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
Prior work retrieves skills from a bank by semantic relevance, then uses them as inference-time patches or for model distillation.
The individual utility of each skill, however, is largely neglected.
We first show that, in on-policy distillation where skill-conditioned policies serve as teachers, fewer than 25% of retrieved skills provide useful distillation signals.
We then propose SGUID, a method for selecting a compact subset of skills for distillation.
SGUID retains a skill only if it consistently yields effective learning signals during training.
The selected skills are then distilled to produce a better model.
Our results show that not all skills are worth distilling.
Across four models from the Olmo and Qwen families, distilling 6 selected skills matches or exceeds full-bank distillation in mean avg@12 on three of the four models, and on all four after a second round that distills 3 newly selected skills, while the full banks are up to 11x larger.
Importantly, SGUID supports stable model-skill co-evolution:
- After a distillation round, a new candidate bank is curated from the updated model's rollouts.
- SGUID selects which skills to internalize next.
In the second round, this loop selects 3 new skills and improves Qwen3-8B from 64.3% to 66.3%.
The selection step is essential for stability:
- On Qwen3-4B, naively updating the model with unfiltered skills degrades performance, including a 0.3 percentage point drop on HMMT25.
- Whereas SGUID improves HMMT25 by 0.5 points after the first round and 1.1 points after the second.
These results identify skill selection as the key mechanism for stable model-skill co-evolution.