The repetition count that works best for a small language model may not remain best at a larger scale.
We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction.
On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens.
We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale.
On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M.
This supports candidate retention as a practical alternative to exact point prediction.
We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms.
A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.