首页 > AI前沿 > CERO: Where and When to Allocate Rollouts for RL Post-Training

CERO: Where and When to Allocate Rollouts for RL Post-Training

arXiv机器学习 2026-10-07 16:40 4 阅读 查看原文

Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts.

We instead study how to coordinate a finite rollout budget over the entire training horizon.

Methodology

We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing.

In our experiments, each admitted prompt receives a fixed-size response group.

CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round.

Technical Details

A compact Fenchel representation linearizes the dependence on cumulative exposure,

while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations.

We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks,

with explicit terms for proxy discrepancy and rate variation.

Results

Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks.

Analysis

Mechanistic analyses link CERO's prompt choices to within-group reward contrast,

while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.