How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lu, Ke-Han, Fu, Szu-Wei, Yang, Chao-Han Huck, Chen, Zhehuai, Huang, Sung-Feng, Yang, Chih-Kai, Lin, Yi-Cheng, Hsiao, Chi-Yuan, Ren, Wenze, Hu, En-Pei, Huang, Yu-Han, Cheng, An-Yu, Chiang, Cheng-Han, Tsao, Yu, Wang, Yu-Chiang Frank, Lee, Hung-yi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911530568646656
author Lu, Ke-Han
Fu, Szu-Wei
Yang, Chao-Han Huck
Chen, Zhehuai
Huang, Sung-Feng
Yang, Chih-Kai
Lin, Yi-Cheng
Hsiao, Chi-Yuan
Ren, Wenze
Hu, En-Pei
Huang, Yu-Han
Cheng, An-Yu
Chiang, Cheng-Han
Tsao, Yu
Wang, Yu-Chiang Frank
Lee, Hung-yi
author_facet Lu, Ke-Han
Fu, Szu-Wei
Yang, Chao-Han Huck
Chen, Zhehuai
Huang, Sung-Feng
Yang, Chih-Kai
Lin, Yi-Cheng
Hsiao, Chi-Yuan
Ren, Wenze
Hu, En-Pei
Huang, Yu-Han
Cheng, An-Yu
Chiang, Cheng-Han
Tsao, Yu
Wang, Yu-Chiang Frank
Lee, Hung-yi
contents Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
Lu, Ke-Han
Fu, Szu-Wei
Yang, Chao-Han Huck
Chen, Zhehuai
Huang, Sung-Feng
Yang, Chih-Kai
Lin, Yi-Cheng
Hsiao, Chi-Yuan
Ren, Wenze
Hu, En-Pei
Huang, Yu-Han
Cheng, An-Yu
Chiang, Cheng-Han
Tsao, Yu
Wang, Yu-Chiang Frank
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Sound
Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.
title How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2603.19195