LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public.
We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator.
Derivation of Sensitivity Curve
For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation.
The relevant correlation must match the score used for ranking and the sampling model:
- In a controlled family, pooled item correlation is 0.90.
- Under item resampling and 0.46 when MMLU subjects are resampled.
- When MMLU subjects are resampled, composite-score correlation is 0.92.
Item-Based Audit
An item-based audit of 394 adjacent-rank claims on the Open LLM Leaderboard finds that 391 lack statistical support even before accounting for selection.
Among claims that pass the uncorrected test, certification can depend on assumptions about the hidden family's correlation.
Resulting Curves
The resulting curves make these assumptions explicit without estimating the unobserved search size.