How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911530568646656 |
|---|---|
| author | Lu, Ke-Han Fu, Szu-Wei Yang, Chao-Han Huck Chen, Zhehuai Huang, Sung-Feng Yang, Chih-Kai Lin, Yi-Cheng Hsiao, Chi-Yuan Ren, Wenze Hu, En-Pei Huang, Yu-Han Cheng, An-Yu Chiang, Cheng-Han Tsao, Yu Wang, Yu-Chiang Frank Lee, Hung-yi |
| author_facet | Lu, Ke-Han Fu, Szu-Wei Yang, Chao-Han Huck Chen, Zhehuai Huang, Sung-Feng Yang, Chih-Kai Lin, Yi-Cheng Hsiao, Chi-Yuan Ren, Wenze Hu, En-Pei Huang, Yu-Han Cheng, An-Yu Chiang, Cheng-Han Tsao, Yu Wang, Yu-Chiang Frank Lee, Hung-yi |
| contents | Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_19195 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation Lu, Ke-Han Fu, Szu-Wei Yang, Chao-Han Huck Chen, Zhehuai Huang, Sung-Feng Yang, Chih-Kai Lin, Yi-Cheng Hsiao, Chi-Yuan Ren, Wenze Hu, En-Pei Huang, Yu-Han Cheng, An-Yu Chiang, Cheng-Han Tsao, Yu Wang, Yu-Chiang Frank Lee, Hung-yi Audio and Speech Processing Computation and Language Sound Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research. |
| title | How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2603.19195 |