Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Leyi, Fu, Zheyu, Zhai, Yunpeng, Tao, Shuchang, Guan, Sheng, Huang, Shiyu, Zhang, Lingzhe, Liu, Zhaoyang, Ding, Bolin, Henry, Felix, Liu, Aiwei, Wen, Lijie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912611065397248
author Pan, Leyi
Fu, Zheyu
Zhai, Yunpeng
Tao, Shuchang
Guan, Sheng
Huang, Shiyu
Zhang, Lingzhe
Liu, Zhaoyang
Ding, Bolin
Henry, Felix
Liu, Aiwei
Wen, Lijie
author_facet Pan, Leyi
Fu, Zheyu
Zhai, Yunpeng
Tao, Shuchang
Guan, Sheng
Huang, Shiyu
Zhang, Lingzhe
Liu, Zhaoyang
Ding, Bolin
Henry, Felix
Liu, Aiwei
Wen, Lijie
contents The rise of Omni-modal Large Language Models (OLLMs), which integrate visual and auditory processing with text, necessitates robust safety evaluations to mitigate harmful outputs. However, no dedicated benchmarks currently exist for OLLMs, and existing benchmarks fail to assess safety under joint audio-visual inputs or cross-modal consistency. To fill this gap, we introduce Omni-SafetyBench, the first comprehensive parallel benchmark for OLLM safety evaluation, featuring 24 modality variations with 972 samples each, including audio-visual harm cases. Considering OLLMs' comprehension challenges with complex omni-modal inputs and the need for cross-modal consistency evaluation, we propose tailored metrics: a Safety-score based on Conditional Attack Success Rate (C-ASR) and Refusal Rate (C-RR) to account for comprehension failures, and a Cross-Modal Safety Consistency score (CMSC-score) to measure consistency across modalities. Evaluating 6 open-source and 4 closed-source OLLMs reveals critical vulnerabilities: (1) only 3 models achieving over 0.6 in both average Safety-score and CMSC-score; (2) safety defenses weaken with complex inputs, especially audio-visual joints; (3) severe weaknesses persist, with some models scoring as low as 0.14 on specific modalities. Using Omni-SafetyBench, we evaluated existing safety alignment algorithms and identified key challenges in OLLM safety alignment: (1) Inference-time methods are inherently less effective as they cannot alter the model's underlying understanding of safety; (2) Post-training methods struggle with out-of-distribution issues due to the vast modality combinations in OLLMs; and, safety tasks involving audio-visual inputs are more complex, making even in-distribution training data less effective. Our proposed benchmark, metrics and the findings highlight urgent needs for enhanced OLLM safety.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
Pan, Leyi
Fu, Zheyu
Zhai, Yunpeng
Tao, Shuchang
Guan, Sheng
Huang, Shiyu
Zhang, Lingzhe
Liu, Zhaoyang
Ding, Bolin
Henry, Felix
Liu, Aiwei
Wen, Lijie
Computation and Language
68T50
I.2.7
The rise of Omni-modal Large Language Models (OLLMs), which integrate visual and auditory processing with text, necessitates robust safety evaluations to mitigate harmful outputs. However, no dedicated benchmarks currently exist for OLLMs, and existing benchmarks fail to assess safety under joint audio-visual inputs or cross-modal consistency. To fill this gap, we introduce Omni-SafetyBench, the first comprehensive parallel benchmark for OLLM safety evaluation, featuring 24 modality variations with 972 samples each, including audio-visual harm cases. Considering OLLMs' comprehension challenges with complex omni-modal inputs and the need for cross-modal consistency evaluation, we propose tailored metrics: a Safety-score based on Conditional Attack Success Rate (C-ASR) and Refusal Rate (C-RR) to account for comprehension failures, and a Cross-Modal Safety Consistency score (CMSC-score) to measure consistency across modalities. Evaluating 6 open-source and 4 closed-source OLLMs reveals critical vulnerabilities: (1) only 3 models achieving over 0.6 in both average Safety-score and CMSC-score; (2) safety defenses weaken with complex inputs, especially audio-visual joints; (3) severe weaknesses persist, with some models scoring as low as 0.14 on specific modalities. Using Omni-SafetyBench, we evaluated existing safety alignment algorithms and identified key challenges in OLLM safety alignment: (1) Inference-time methods are inherently less effective as they cannot alter the model's underlying understanding of safety; (2) Post-training methods struggle with out-of-distribution issues due to the vast modality combinations in OLLMs; and, safety tasks involving audio-visual inputs are more complex, making even in-distribution training data less effective. Our proposed benchmark, metrics and the findings highlight urgent needs for enhanced OLLM safety.
title Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2508.07173