首页 > AI前沿 > CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

arXiv机器学习 2026-08-28 05:43 5 阅读 查看原文

Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown.

We present CleanScore, a black-box audit using scored outputs only.

Each benchmark question becomes a parent item with one public form and two independently written fresh forms preserving its numbers, facts and answer.

The audit reports an interval for the public-form advantage rather than a verdict, and a private negative-control bank with an explicit transport radius separates exposure from ordinary form mismatch.

A registered controlled-exposure experiment detects planted exposure and stays quiet under fresh-form exposure.

A registered audit of five open models on 200 GSM8K and 200 ARC-Challenge items finds no exposure-consistent advantage, bounding surface-form inflation below five points.

Registered positive controls then bound what such a null can mean.

Leaking an item raises accuracy on paraphrases the model never saw almost as much as on the leaked wording, leaving 52% to 110% of the effect invisible to a paraphrase audit.

On ARC a planted 49-point advantage shows an observable gap of -0.020, and about 20 points survive rewriting stem and options, across four training seeds.

A surface-form null bounds far less than the phrase contamination audit implies.