Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown.
We present CleanScore, a black-box audit using scored outputs only.
Each benchmark question becomes a parent item with one public form and two independently written fresh forms preserving its numbers, facts and answer.
The audit reports an interval for the public-form advantage rather than a verdict, and a private negative-control bank with an explicit transport radius separates exposure from ordinary form mismatch.
A registered controlled-exposure experiment detects planted exposure and stays quiet under fresh-form exposure.
A registered audit of five open models on 200 GSM8K and 200 ARC-Challenge items finds no exposure-consistent advantage, bounding surface-form inflation below five points.
Registered positive controls then bound what such a null can mean.
Leaking an item raises accuracy on paraphrases the model never saw almost as much as on the leaked wording, leaving 52% to 110% of the effect invisible to a paraphrase audit.
On ARC a planted 49-point advantage shows an observable gap of -0.020, and about 20 points survive rewriting stem and options, across four training seeds.
A surface-form null bounds far less than the phrase contamination audit implies.