SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chih-Kai, Ho, Neo, Piao, Yen-Ting, Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909750504980480
author Yang, Chih-Kai
Ho, Neo
Piao, Yen-Ting
Lee, Hung-yi
author_facet Yang, Chih-Kai
Ho, Neo
Piao, Yen-Ting
Lee, Hung-yi
contents Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13237
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Yang, Chih-Kai
Ho, Neo
Piao, Yen-Ting
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Sound
Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research.
title SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2505.13237