SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909750504980480 |
|---|---|
| author | Yang, Chih-Kai Ho, Neo Piao, Yen-Ting Lee, Hung-yi |
| author_facet | Yang, Chih-Kai Ho, Neo Piao, Yen-Ting Lee, Hung-yi |
| contents | Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13237 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information Yang, Chih-Kai Ho, Neo Piao, Yen-Ting Lee, Hung-yi Audio and Speech Processing Computation and Language Sound Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research. |
| title | SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2505.13237 |