首页 > AI前沿 > GRPO Training Dynamics for Small Language Models

GRPO Training Dynamics for Small Language Models

arXiv机器学习 2026-09-30 16:59 7 阅读 查看原文

Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks.

However, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource-constrained environments.

In this work

We present a systematic study of GRPO fine-tuning for SLMs ranging from 1.5B to 7B parameters under a practical single-node 8xA100 compute budget.

Our study spans multiple model families and reasoning domains, including mathematics, coding, and multiple-choice question answering (MCQ) in science.

Across these settings, we analyze how group size affects policy convergence, training stability, and downstream benchmark performance.

We further characterize tensor-level update dynamics during GRPO training and investigate whether the choice of LoRA target modules and layers can improve the performance of GRPO-tuned models.

While our initial GRPO-tuned models outperform their base counterparts on approximately 80% of mathematical benchmark evaluations, they demonstrate limited capability on MCQ and code reasoning tasks.

Guided by our mechanistic evaluations, we refined our LoRA and reward-shaping configurations to improve performance in latter domains.

These findings provide practical guidance for GRPO training for SLMs.