首页 > AI前沿 > KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

arXiv自然语言 2026-05-27 14:08 5 阅读 查看原文

Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

Current MLLM safety evaluation tools, however, suffer from major limitations: 1) English-centric dataset construction, and 2) a focus on generic risks that are not tied to local cultural contexts.

This paper introduces KSAFE-MM

a benchmark for Korean multimodal safety evaluation that covers both general safety risks and culture-specific vulnerabilities.

KSAFE-MM consists of two complementary parts:

  • KSAFE-MM-G evaluates globally shared risks in Korean contexts through linguistic contextualization, which transforms generic safety queries into contextually grounded multimodal samples.

  • In contrast, KSAFE-MM-C targets safety vulnerabilities that are culture-dependent, using localized visual queries drawn from real-world. It pairs these visual queries with jailbreak-style textual queries to cover multimodal safety risks involving cultural visual cues and malicious textual intent.

We evaluate 12 state-of-the-art MLLMs on KSAFE-MM and reveal culturally grounded vulnerabilities that translation-based evaluation fails to capture.

Notably, jailbreaking strategies substantially amplify attack success rates, with ProgramExecution yielding up to 74.2% ASR compared to 13.4% for standard queries.

Furthermore, we identify a systematic trade-off between safety and over-refusal, where models achieving low ASR tend to exhibit excessive refusal behavior on benign queries.

These findings highlight the urgent need for culturally grounded safety evaluation beyond English-centric benchmarks.