首页 > AI前沿 > Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

arXiv自然语言 2026-09-23 08:28 5 阅读 查看原文

Evidence that evaluation material entered training does not reveal how much it affected evaluation.

This distinction leaves a contaminated benchmark score difficult to interpret:

provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact.

We present LeakScale, an interventional framework for estimating this missing quantity.

LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-derivable from the public task, controls access to that information, and estimates the resulting control-adjusted change in executable accuracy.

Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improves accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points.

These findings separate two empirical questions that are often conflated:

  • whether benchmark contact occurred
  • how strongly a reported score depends on it

LeakScale makes the latter directly measurable.