Benchmarks for large language models (LLMs)
Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels.
But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem.
Evaluate this capability of information discovery
To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question.
CruxBench enjoys a rare combination of three key properties:
- contamination-resistant by construction, since ground truth is generated by future world events;
- open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer;
- grounded, with informativeness measured against quantified changes in real-world beliefs.
We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes' usefulness for answering target questions.
However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.