首页 > AI前沿 > CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

arXiv自然语言 2026-10-06 01:59 7 阅读 查看原文

Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation.

We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema.

The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields.

Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit.

We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass.

The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable.

We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.