首页 > AI前沿 > PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

arXiv自然语言 2026-09-23 13:52 5 阅读 查看原文

Compact vision-language models (VLMs) now power a growing share of multimodal applications.

The benchmarks used to compare them, however, inherit a frontier-centric design:

  • each model is reduced to a single accuracy number,
  • narrowing the inter-model gap on saturated suites,
  • and pressing models into low-score bands on harder ones.

We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks.

Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away.

Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy.

We will release the full pipeline, prompts, and per-item annotations.