首页 > AI前沿 > ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

arXiv机器学习 2026-09-30 04:57 6 阅读 查看原文

Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies.

This leads to overly optimistic generalization estimates.

Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets.

Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes.

We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting.

ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling.

ReLaG is open source and can be installed with pip install relag.