首页 > AI前沿 > MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

arXiv自然语言 2026-10-08 21:41 4 阅读 查看原文

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute.

Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up.

Scaling RL Compute

We scale RL compute along three dimensions:

  • Larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M;
  • More diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses;
  • More grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions.

Stability and Infrastructure

To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking.

We further build infrastructure for mixed-task agentic RL, including:

  • A unified trajectory representation;
  • High-concurrency multi-framework rollout;
  • Decoupled control and data planes;
  • Training-inference consistency.

Open Source Contributions

We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.