A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Kaiwen, Zhou, Zhenhong, Wang, Leo, Lin, Liang, Xiao, Yang, Shao, Tianyu, Zhang, Yuanhe, Li, Yuxuan, Yu, Miao, Lyu, Kailin, Zhang, Jiaming, Liu, Dongrui, Sun, Li, Wu, Yueming, Li, Kai, Dang, Ting, Jia, Xiaojun, Das, Rohan Kumar, Li, Xinfeng, Liang, Siyuan, Wang, Qiufeng, Ma, Xingjun, Chen, Jing, Wang, Kun, Dong, Junhao, Zou, Deqing, Cheng, Yu, Hu, Xia, Zeng, Zhigang, Su, Sen, Liu, Yang, Jiang, Yu-Gang, Yu, Philip S., Ong, Yew-Soon
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910237184753664
author Luo, Kaiwen
Zhou, Zhenhong
Wang, Leo
Lin, Liang
Xiao, Yang
Shao, Tianyu
Zhang, Yuanhe
Li, Yuxuan
Yu, Miao
Lyu, Kailin
Zhang, Jiaming
Liu, Dongrui
Sun, Li
Wu, Yueming
Li, Kai
Dang, Ting
Jia, Xiaojun
Das, Rohan Kumar
Li, Xinfeng
Liang, Siyuan
Wang, Qiufeng
Ma, Xingjun
Chen, Jing
Wang, Kun
Dong, Junhao
Zou, Deqing
Cheng, Yu
Hu, Xia
Zeng, Zhigang
Su, Sen
Liu, Yang
Jiang, Yu-Gang
Yu, Philip S.
Ong, Yew-Soon
author_facet Luo, Kaiwen
Zhou, Zhenhong
Wang, Leo
Lin, Liang
Xiao, Yang
Shao, Tianyu
Zhang, Yuanhe
Li, Yuxuan
Yu, Miao
Lyu, Kailin
Zhang, Jiaming
Liu, Dongrui
Sun, Li
Wu, Yueming
Li, Kai
Dang, Ting
Jia, Xiaojun
Das, Rohan Kumar
Li, Xinfeng
Liang, Siyuan
Wang, Qiufeng
Ma, Xingjun
Chen, Jing
Wang, Kun
Dong, Junhao
Zou, Deqing
Cheng, Yu
Hu, Xia
Zeng, Zhigang
Su, Sen
Liu, Yang
Jiang, Yu-Gang
Yu, Philip S.
Ong, Yew-Soon
contents The foundational capabilities established by Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs), within which Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This survey provides a comprehensive investigation into the endogenous mechanisms of LALMs, detailing the architectural innovations and alignment algorithms that facilitate emergent reasoning. Specifically, we analyze how the transition to unified end-to-end frameworks and the integration of continuous acoustic signals inherently expand the attack surface. To rigorously evaluate the risks within these paradigms, we establish a comprehensive taxonomy of trustworthiness, categorizing critical vulnerabilities such as cross-modal jailbreaking, latent acoustic backdoors, and biometric privacy leakage. We review the state-of-the-art through six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. The profound imbalance between a mature offensive landscape and underdeveloped defenses further validates the critical trustworthiness gaps and multidimensional risks facing audio-centric intelligence. Finally, we propose a strategic roadmap advocating for "Defense-in-Depth" architectures, causal auditory world modeling, and intrinsic representation engineering to bridge the gap between empirical performance and intrinsically trustworthy audio intelligence. Our project has been uploaded to GitHub https://github.com/Kwwwww74/Awesome-Trustworthy-AudioLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20266
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
Luo, Kaiwen
Zhou, Zhenhong
Wang, Leo
Lin, Liang
Xiao, Yang
Shao, Tianyu
Zhang, Yuanhe
Li, Yuxuan
Yu, Miao
Lyu, Kailin
Zhang, Jiaming
Liu, Dongrui
Sun, Li
Wu, Yueming
Li, Kai
Dang, Ting
Jia, Xiaojun
Das, Rohan Kumar
Li, Xinfeng
Liang, Siyuan
Wang, Qiufeng
Ma, Xingjun
Chen, Jing
Wang, Kun
Dong, Junhao
Zou, Deqing
Cheng, Yu
Hu, Xia
Zeng, Zhigang
Su, Sen
Liu, Yang
Jiang, Yu-Gang
Yu, Philip S.
Ong, Yew-Soon
Sound
The foundational capabilities established by Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs), within which Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This survey provides a comprehensive investigation into the endogenous mechanisms of LALMs, detailing the architectural innovations and alignment algorithms that facilitate emergent reasoning. Specifically, we analyze how the transition to unified end-to-end frameworks and the integration of continuous acoustic signals inherently expand the attack surface. To rigorously evaluate the risks within these paradigms, we establish a comprehensive taxonomy of trustworthiness, categorizing critical vulnerabilities such as cross-modal jailbreaking, latent acoustic backdoors, and biometric privacy leakage. We review the state-of-the-art through six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. The profound imbalance between a mature offensive landscape and underdeveloped defenses further validates the critical trustworthiness gaps and multidimensional risks facing audio-centric intelligence. Finally, we propose a strategic roadmap advocating for "Defense-in-Depth" architectures, causal auditory world modeling, and intrinsic representation engineering to bridge the gap between empirical performance and intrinsically trustworthy audio intelligence. Our project has been uploaded to GitHub https://github.com/Kwwwww74/Awesome-Trustworthy-AudioLLMs.
title A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
topic Sound
url https://arxiv.org/abs/2605.20266