首页 > AI前沿 > CruxBench: A Benchmark of Information Discovery

CruxBench: A Benchmark of Information Discovery

arXiv自然语言 2026-09-27 07:28 5 阅读 查看原文

Benchmarks for large language models (LLMs)

Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels.

But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem.

Evaluate this capability of information discovery

To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question.

CruxBench enjoys a rare combination of three key properties:

  • contamination-resistant by construction, since ground truth is generated by future world events;
  • open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer;
  • grounded, with informativeness measured against quantified changes in real-world beliefs.

We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes' usefulness for answering target questions.

However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.