首页 > AI前沿 > When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection

When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection

arXiv机器学习 2026-08-27 15:43 4 阅读 查看原文

A released table is often treated as an i.i.d. sample, although its repeated rows may encode business frequency, repeated entities, joins, resampling, or extraction errors.

We show that this ambiguity creates a hidden measurement layer with three consequences: feature-identical rows impose an attained evaluation ceiling, row-weighted AUROC is sensitive to replication, and row-trained detectors learn a multiplicity-size-biased law.

An exact-row audit of all 690 OddBench datasets finds train-test overlap in 355, feature-identical label conflict in 147, and a test anomaly identical to a training normal in 137.

Switching from row to support weighting changes AUROC by at least 0.05 on 50-61 datasets across four classical detector geometries.

We introduce SCOUT (Support-Count Orthogonalized Unsupervised Testing), a factorized anomaly detector that separates replication-invariant support evidence from exposure-aware count evidence.

Factorwise split-conformal calibration yields marginal false-positive-rate control, while the support channel is exactly invariant to arbitrary positive row replication.

On 686 OddBench datasets and five seeds, support-only SCOUT is non-inferior to row-wise Isolation Forest in raw AUROC and improves replication-invariant AUROC.

External normal-support evaluations track nominal false-positive levels, and four backbones remain exactly unchanged under controlled replication.

Semi-synthetic interventions show that conditional count modeling helps materially only under strong rate heterogeneity.

These results specify when multiplicity should be treated as signal, nuisance, or uninterpretable without additional information.