首页 > AI前沿 > ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual Density

ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual Density

arXiv机器学习 2026-09-30 10:13 5 阅读 查看原文

Multimodal large language models (MLLMs) offer a promising approach for recovering numerical data from scientific charts, but their ability to reconstruct chart data from visually dense figures remains poorly understood.

Existing chart understanding benchmarks primarily evaluate question answering or chart-level reasoning and provide limited support for evaluating structured numerical reconstruction from scientific figures.

We introduce \textbf{ChartDensity-Bench}, a benchmark for evaluating MLLMs on structured numerical data reconstruction from compound chart figures under controlled visual density.

Built from charts paired with source-level ground-truth data, ChartDensity-Bench systematically varies the number of simultaneously presented charts ($k\in{1,3,6,9}$), enabling controlled evaluation of density-induced degradation.

We further propose a multi-dimensional evaluation framework covering structural reliability, reconstruction completeness, parseability, and numerical fidelity.

Experiments on five recent MLLMs show that numerical reconstruction generally degrades as visual density increases, while the magnitude of degradation varies substantially across models.

Chart-level paired comparisons further show that the same source chart can incur higher reconstruction error when embedded in denser visual contexts.

These findings highlight visual density as an important and previously underexplored factor in MLLM chart data reconstruction and provide a systematic benchmark for evaluating model robustness in this setting.