首页 > AI前沿 > Less Uniform Discrete Diffusion is More Powerful and Scalable

Less Uniform Discrete Diffusion is More Powerful and Scalable

arXiv自然语言 2026-09-21 21:25 6 阅读 查看原文

Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging.

We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling.

To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework.

Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling.

Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation.

We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning.

It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.