Evidence that evaluation material entered training does not reveal how much it affected evaluation.
This distinction leaves a contaminated benchmark score difficult to interpret:
provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact.
We present LeakScale, an interventional framework for estimating this missing quantity.
LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-derivable from the public task, controls access to that information, and estimates the resulting control-adjusted change in executable accuracy.
Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improves accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points.
These findings separate two empirical questions that are often conflated:
- whether benchmark contact occurred
- how strongly a reported score depends on it
LeakScale makes the latter directly measurable.