Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Zirui, Jiang, Qian, Cui, Mingxuan, Li, Mingzhe, Gao, Lang, Zhang, Zeyu, Xu, Zixiang, Wang, Yanbo, Wang, Chenxi, Ouyang, Guangxian, Chen, Zhenhao, Chen, Xiuying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910959481651200
author Song, Zirui
Jiang, Qian
Cui, Mingxuan
Li, Mingzhe
Gao, Lang
Zhang, Zeyu
Xu, Zixiang
Wang, Yanbo
Wang, Chenxi
Ouyang, Guangxian
Chen, Zhenhao
Chen, Xiuying
author_facet Song, Zirui
Jiang, Qian
Cui, Mingxuan
Li, Mingzhe
Gao, Lang
Zhang, Zeyu
Xu, Zixiang
Wang, Yanbo
Wang, Chenxi
Ouyang, Guangxian
Chen, Zhenhao
Chen, Xiuying
contents The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LAM safety especially against jailbreak attacks, which are challenging due to the temporal and semantic nature of speech. To bridge this gap, we introduce AJailBench, the first benchmark specifically designed to evaluate jailbreak vulnerabilities in LAMs. We begin by constructing AJailBench-Base, a dataset of 1,495 adversarial audio prompts spanning 10 policy-violating categories, converted from textual jailbreak attacks using realistic text to speech synthesis. Using this dataset, we evaluate several state-of-the-art LAMs and reveal that none exhibit consistent robustness across attacks. To further strengthen jailbreak testing and simulate more realistic attack conditions, we propose a method to generate dynamic adversarial variants. Our Audio Perturbation Toolkit (APT) applies targeted distortions across time, frequency, and amplitude domains. To preserve the original jailbreak intent, we enforce a semantic consistency constraint and employ Bayesian optimization to efficiently search for perturbations that are both subtle and highly effective. This results in AJailBench-APT, an extended dataset of optimized adversarial audio samples. Our findings demonstrate that even small, semantically preserved perturbations can significantly reduce the safety performance of leading LAMs, underscoring the need for more robust and semantically aware defense mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
Song, Zirui
Jiang, Qian
Cui, Mingxuan
Li, Mingzhe
Gao, Lang
Zhang, Zeyu
Xu, Zixiang
Wang, Yanbo
Wang, Chenxi
Ouyang, Guangxian
Chen, Zhenhao
Chen, Xiuying
Sound
Artificial Intelligence
Audio and Speech Processing
The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LAM safety especially against jailbreak attacks, which are challenging due to the temporal and semantic nature of speech. To bridge this gap, we introduce AJailBench, the first benchmark specifically designed to evaluate jailbreak vulnerabilities in LAMs. We begin by constructing AJailBench-Base, a dataset of 1,495 adversarial audio prompts spanning 10 policy-violating categories, converted from textual jailbreak attacks using realistic text to speech synthesis. Using this dataset, we evaluate several state-of-the-art LAMs and reveal that none exhibit consistent robustness across attacks. To further strengthen jailbreak testing and simulate more realistic attack conditions, we propose a method to generate dynamic adversarial variants. Our Audio Perturbation Toolkit (APT) applies targeted distortions across time, frequency, and amplitude domains. To preserve the original jailbreak intent, we enforce a semantic consistency constraint and employ Bayesian optimization to efficiently search for perturbations that are both subtle and highly effective. This results in AJailBench-APT, an extended dataset of optimized adversarial audio samples. Our findings demonstrate that even small, semantically preserved perturbations can significantly reduce the safety performance of leading LAMs, underscoring the need for more robust and semantically aware defense mechanisms.
title Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.15406