Compact vision-language models (VLMs) now power a growing share of multimodal applications.
The benchmarks used to compare them, however, inherit a frontier-centric design:
- each model is reduced to a single accuracy number,
- narrowing the inter-model gap on saturated suites,
- and pressing models into low-score bands on harder ones.
We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks.
Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away.
Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy.
We will release the full pipeline, prompts, and per-item annotations.