首页 > AI前沿 > Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

arXiv自然语言 2026-09-22 17:31 6 阅读 查看原文

Diffusion Large Language Models (dLLMs) have emerged as an efficient alternative to autoregressive models, yet aligning them via Reinforcement Learning (RL) requires likelihood surrogates estimated from masked reconstruction subproblems under a small Monte Carlo budget per rollout.

Existing methods construct these subproblems by uniform random masking, leaving open the question of which subproblems to prioritize.

We identify a systematic upstream/downstream structure in dLLM rollouts. Some tokens, when revealed, trigger large confidence changes in nearby undecoded positions; we call them upstream. Others induce only small local changes and are therefore downstream.

We find masking downstream tokens yields substantially better-posed subproblems than masking upstream tokens, a phenomenon we term subproblem difficulty asymmetry.

Based on the observation, we propose Informed Masking (IM), which derives a per-token priority score from the denoising trajectory at zero extra inference cost and biases mask sampling toward downstream tokens.

IM is plug-and-play: when plugged into three state-of-the-art dLLM RL methods on LLaDA-8B-Instruct, it delivers up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks with improved training stability.