Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Bo-Han, Liu, Chien-Feng, Liang, Yu-Hsuan Li, Yang, Chih-Kai, Fu, Szu-Wei, Chen, Zhehuai, Lu, Ke-Han, Huang, Sung-Feng, Yang, Chao-Han Huck, Wang, Yu-Chiang Frank, Chen, Yun-Nung, Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918163869859840
author Feng, Bo-Han
Liu, Chien-Feng
Liang, Yu-Hsuan Li
Yang, Chih-Kai
Fu, Szu-Wei
Chen, Zhehuai
Lu, Ke-Han
Huang, Sung-Feng
Yang, Chao-Han Huck
Wang, Yu-Chiang Frank
Chen, Yun-Nung
Lee, Hung-yi
author_facet Feng, Bo-Han
Liu, Chien-Feng
Liang, Yu-Hsuan Li
Yang, Chih-Kai
Fu, Szu-Wei
Chen, Zhehuai
Lu, Ke-Han
Huang, Sung-Feng
Yang, Chao-Han Huck
Wang, Yu-Chiang Frank
Chen, Yun-Nung
Lee, Hung-yi
contents Large audio-language models (LALMs) extend text-based LLMs with auditory understanding, offering new opportunities for multimodal applications. While their perception, reasoning, and task performance have been widely studied, their safety alignment under paralinguistic variation remains underexplored. This work systematically investigates the role of speaker emotion. We construct a dataset of malicious speech instructions expressed across multiple emotions and intensities, and evaluate several state-of-the-art LALMs. Our results reveal substantial safety inconsistencies: different emotions elicit varying levels of unsafe responses, and the effect of intensity is non-monotonic, with medium expressions often posing the greatest risk. These findings highlight an overlooked vulnerability in LALMs and call for alignment strategies explicitly designed to ensure robustness under emotional variation, a prerequisite for trustworthy deployment in real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16893
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
Feng, Bo-Han
Liu, Chien-Feng
Liang, Yu-Hsuan Li
Yang, Chih-Kai
Fu, Szu-Wei
Chen, Zhehuai
Lu, Ke-Han
Huang, Sung-Feng
Yang, Chao-Han Huck
Wang, Yu-Chiang Frank
Chen, Yun-Nung
Lee, Hung-yi
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Large audio-language models (LALMs) extend text-based LLMs with auditory understanding, offering new opportunities for multimodal applications. While their perception, reasoning, and task performance have been widely studied, their safety alignment under paralinguistic variation remains underexplored. This work systematically investigates the role of speaker emotion. We construct a dataset of malicious speech instructions expressed across multiple emotions and intensities, and evaluate several state-of-the-art LALMs. Our results reveal substantial safety inconsistencies: different emotions elicit varying levels of unsafe responses, and the effect of intensity is non-monotonic, with medium expressions often posing the greatest risk. These findings highlight an overlooked vulnerability in LALMs and call for alignment strategies explicitly designed to ensure robustness under emotional variation, a prerequisite for trustworthy deployment in real-world settings.
title Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2510.16893