Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible.
We study condition-anchored generative distillation (CAGD): retain a small set of old prompts, use a frozen previous model to reconstruct completions and generation states, and match its predictive distributions while learning the next task.
The formulation separates three roles that ordinary replay conflates: conditions select the behavior to protect, teacher generations locate relevant states, and soft targets specify how predictions may change.
Autoregressive Language Generation
For autoregressive language generation, teacher-rollout distillation admits an exact chain-rule decomposition of sequence divergence.
Masked-Diffusion Language Modeling
For masked-diffusion language modeling, our implementation directly controls local denoising drift on teacher-generated completions.
Continual Adaptation Results
In continual adaptation of a 219M masked diffusion language model, CAGD reduces four-task final held-out loss from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse.
The same soft targets lower final average loss by 0.055 over hard replay when teacher-generated support is held identical.
The direction persists on fresh facts and natural instructions across SMDM and Qwen3.
Adaptation on GSM8K
On GSM8K, Qwen adaptation preserves answer-format compliance, but exact-match retention is seed-mixed at 0.6B and worsens at 1.7B.
These results support condition-anchored functional preservation as a common design principle across the tested language-generation objectives.