MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhongxi, Lin, Yueqian, Zhang, Jingyang, Li, Hai Helen, Chen, Yiran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918367313526784
author Wang, Zhongxi
Lin, Yueqian
Zhang, Jingyang
Li, Hai Helen
Chen, Yiran
author_facet Wang, Zhongxi
Lin, Yueqian
Zhang, Jingyang
Li, Hai Helen
Chen, Yiran
contents Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, run-centric platform that integrates automatic cross-modal payload generation, three multi-turn attack algorithms (Crescendo, PAIR, Violent Durian), provider-agnostic model routing, and an LLM judge with a five-level safety taxonomy into a single browser-based system. A dual-metric framework distinguishes hard Attack Success Rate (Compliance only) from soft ASR (including Partial Compliance), capturing partial information leakage that binary metrics miss. To probe whether alignment generalizes across modality boundaries, we introduce Inter-Turn Modality Switching (ITMS), which augments multi-turn attacks with per-turn modality rotation. Experiments across six multimodal LLMs from four providers show that multi-turn strategies can achieve up to 90-100% ASR against models with near-perfect single-turn refusal. ITMS does not uniformly raise final ASR on already-saturated baselines, but accelerates convergence by destabilizing early-turn defenses, and ablation reveals that the direction of modality effects is model-family-specific rather than universal, underscoring the need for provider-aware cross-modal safety testing.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02482
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
Wang, Zhongxi
Lin, Yueqian
Zhang, Jingyang
Li, Hai Helen
Chen, Yiran
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, run-centric platform that integrates automatic cross-modal payload generation, three multi-turn attack algorithms (Crescendo, PAIR, Violent Durian), provider-agnostic model routing, and an LLM judge with a five-level safety taxonomy into a single browser-based system. A dual-metric framework distinguishes hard Attack Success Rate (Compliance only) from soft ASR (including Partial Compliance), capturing partial information leakage that binary metrics miss. To probe whether alignment generalizes across modality boundaries, we introduce Inter-Turn Modality Switching (ITMS), which augments multi-turn attacks with per-turn modality rotation. Experiments across six multimodal LLMs from four providers show that multi-turn strategies can achieve up to 90-100% ASR against models with near-perfect single-turn refusal. ITMS does not uniformly raise final ASR on already-saturated baselines, but accelerates convergence by destabilizing early-turn defenses, and ablation reveals that the direction of modality effects is model-family-specific rather than universal, underscoring the need for provider-aware cross-modal safety testing.
title MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2603.02482