首页 > AI前沿 > MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

arXiv自然语言 2026-08-08 10:49 9 阅读 查看原文

While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters.

Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation.

To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels.

Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors.

Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems.

Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning.

These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.