AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Zhe, Cai, Yujun, Yao, Jiayu, Wang, Yiwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915747715874816
author Sun, Zhe
Cai, Yujun
Yao, Jiayu
Wang, Yiwei
author_facet Sun, Zhe
Cai, Yujun
Yao, Jiayu
Wang, Yiwei
contents Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can perceive spatial dynamics, particularly the motion of sound sources, remains unclear. In this work, we uncover a systematic motion perception deficit in current ALLMs. To investigate this issue, we introduce AudioMotionBench, the first benchmark explicitly designed to evaluate auditory motion understanding. AudioMotionBench introduces a controlled question-answering benchmark designed to evaluate whether Audio-Language Models (LALMs) can infer the direction and trajectory of moving sound sources from binaural audio. Comprehensive quantitative and qualitative analyses reveal that current models struggle to reliably recognize motion cues or distinguish directional patterns. The average accuracy remains below 50\%, underscoring a fundamental limitation in auditory spatial reasoning. Our study highlights a fundamental gap between human and model auditory spatial reasoning, providing both a diagnostic tool and new insight for enhancing spatial cognition in future Audio-Language Models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
Sun, Zhe
Cai, Yujun
Yao, Jiayu
Wang, Yiwei
Sound
Artificial Intelligence
Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can perceive spatial dynamics, particularly the motion of sound sources, remains unclear. In this work, we uncover a systematic motion perception deficit in current ALLMs. To investigate this issue, we introduce AudioMotionBench, the first benchmark explicitly designed to evaluate auditory motion understanding. AudioMotionBench introduces a controlled question-answering benchmark designed to evaluate whether Audio-Language Models (LALMs) can infer the direction and trajectory of moving sound sources from binaural audio. Comprehensive quantitative and qualitative analyses reveal that current models struggle to reliably recognize motion cues or distinguish directional patterns. The average accuracy remains below 50\%, underscoring a fundamental limitation in auditory spatial reasoning. Our study highlights a fundamental gap between human and model auditory spatial reasoning, providing both a diagnostic tool and new insight for enhancing spatial cognition in future Audio-Language Models.
title AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2511.13273