首页 > AI前沿 > Diversifying RLVR Rollouts via First-Token Exploration

Diversifying RLVR Rollouts via First-Token Exploration

arXiv自然语言 2026-10-01 12:00 7 阅读 查看原文

Reinforcement learning with verifiable rewards (RLVR) trains reasoning models without labeled trajectories, using groups of verifier-scored rollouts to explore alternative reasoning paths.

Limited rollout diversity is a central bottleneck, typically addressed through adjustments to temperature, prefixes, or rollout selection.

We identify the first token of the response as a structurally distinct target for diversification, largely overlooked in prior work.

We find that the first-token distribution is sharply concentrated and only weakly related to downstream correctness, as lower-probability candidates can yield similarly accurate responses.

Diversifying the first token can therefore broaden the reasoning paths explored within each rollout group with little loss in response quality.

Motivated by this observation, we introduce REFT (Rollout Exploration with First-Token Diversification), a lightweight modification to RLVR.

REFT samples first tokens uniformly from the policy's top-$N$ candidates and allocates rollouts evenly across the sampled tokens, leaving the rest of the pipeline unchanged.

We evaluate REFT on eight models spanning multiple architectures and sizes (0.5B-14B), with mathematical reasoning and code-generation tasks under GRPO and DAPO.

Across these settings, REFT consistently improves Pass@1, Pass@8, and Pass@64.

It also outperforms competing diversification methods at every evaluated budget, incurring the lowest rollout cost.