首页 > AI前沿 > Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

arXiv自然语言 2026-09-29 13:20 5 阅读 查看原文

Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels.

Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly.

However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group.

Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization.

To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages.

Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization.

GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.