首页 > AI前沿 > Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining

Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining

arXiv自然语言 2026-10-04 19:16 5 阅读 查看原文

The repetition count that works best for a small language model may not remain best at a larger scale.

We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction.

On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens.

We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale.

On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M.

This supports candidate retention as a practical alternative to exact point prediction.

We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms.

A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.