JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Zifan, Liu, Yule, Sun, Zhen, Li, Mingchen, Luo, Zeren, Zheng, Jingyi, Dong, Wenhan, He, Xinlei, Wang, Xuechao, Xue, Yingjie, Xu, Shengmin, Huang, Xinyi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912933132369920
author Peng, Zifan
Liu, Yule
Sun, Zhen
Li, Mingchen
Luo, Zeren
Zheng, Jingyi
Dong, Wenhan
He, Xinlei
Wang, Xuechao
Xue, Yingjie
Xu, Shengmin
Huang, Xinyi
author_facet Peng, Zifan
Liu, Yule
Sun, Zhen
Li, Mingchen
Luo, Zeren
Zheng, Jingyi
Dong, Wenhan
He, Xinlei
Wang, Xuechao
Xue, Yingjie
Xu, Shengmin
Huang, Xinyi
contents Large Audio Language Models (LALMs) have made significant progress. While increasingly deployed in real-world applications, LALMs face growing safety risks from jailbreak attacks that bypass safety alignment. However, there remains a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare jailbreak attacks against them. To address this gap, we introduce JALMBench, a comprehensive benchmark that assesses LALM safety against jailbreak attacks, comprising 11,316 text samples and 245,355 audio samples (>1,000 hours). JALMBench supports 12 mainstream LALMs, 8 attack methods (4 text-transferred and 4 audio-originated), and 5 defenses. We conduct in-depth analysis on attack efficiency, topic sensitivity, voice diversity, and model architecture. Additionally, we explore mitigation strategies for the attacks at both the prompt and response levels. Our systematic evaluation reveals that LALMs' safety is strongly influenced by modality and architectural choices: text-based safety alignment can partially transfer to audio inputs, and interleaved audio-text strategies enable more robust cross-modal generalization. Existing general-purpose moderation methods only slightly improve security, highlighting the need for defense methods specifically designed for LALMs. We hope our work can shed light on the design principles for building more robust LALMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17568
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
Peng, Zifan
Liu, Yule
Sun, Zhen
Li, Mingchen
Luo, Zeren
Zheng, Jingyi
Dong, Wenhan
He, Xinlei
Wang, Xuechao
Xue, Yingjie
Xu, Shengmin
Huang, Xinyi
Cryptography and Security
Artificial Intelligence
Sound
Audio and Speech Processing
Large Audio Language Models (LALMs) have made significant progress. While increasingly deployed in real-world applications, LALMs face growing safety risks from jailbreak attacks that bypass safety alignment. However, there remains a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare jailbreak attacks against them. To address this gap, we introduce JALMBench, a comprehensive benchmark that assesses LALM safety against jailbreak attacks, comprising 11,316 text samples and 245,355 audio samples (>1,000 hours). JALMBench supports 12 mainstream LALMs, 8 attack methods (4 text-transferred and 4 audio-originated), and 5 defenses. We conduct in-depth analysis on attack efficiency, topic sensitivity, voice diversity, and model architecture. Additionally, we explore mitigation strategies for the attacks at both the prompt and response levels. Our systematic evaluation reveals that LALMs' safety is strongly influenced by modality and architectural choices: text-based safety alignment can partially transfer to audio inputs, and interleaved audio-text strategies enable more robust cross-modal generalization. Existing general-purpose moderation methods only slightly improve security, highlighting the need for defense methods specifically designed for LALMs. We hope our work can shed light on the design principles for building more robust LALMs.
title JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
topic Cryptography and Security
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17568