首页 > AI前沿 > UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

arXiv机器学习 2026-09-24 01:01 5 阅读 查看原文

Benchmarking is an increasingly critical part of research in machine learning and the domains where it is applied, including healthcare.

Yet, despite the steady development of new ultrasound foundation models in recent years, the development of well-designed benchmarks to evaluate them has lagged behind.

This deficiency has led to fragmented and inconsistent evaluations of competing models, making it difficult to measure progress.

To address this issue

we introduce UltraBench 2, a comprehensive benchmark with wide anatomical and task coverage, and a focus on standardization, reproducibility, and ease-of-use.

Using this benchmark, we compare existing vision foundation models for ultrasound image analysis.

Our analyses demonstrate that ultrasound-specific pretraining still leads on classification, but that state-of-the-art general-purpose models have drawn level on segmentation.