Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jieyi, Niu, Yazhe, Xu, Dexuan, Wei, Zhongyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911598047657984
author Wang, Jieyi
Niu, Yazhe
Xu, Dexuan
Wei, Zhongyu
author_facet Wang, Jieyi
Niu, Yazhe
Xu, Dexuan
Wei, Zhongyu
contents Recent Large Audio Language Models have demonstrated impressive capabilities in audio understanding. However, they often suffer from perceptual errors, while reliable audio reasoning is unattainable without first grounding the model's perception in structured auditory scenes. Inspired by Auditory Scene Analysis, we first introduce a Perception-Aware Question Answering (PAQA) dataset. PAQA implements a hierarchical decoupling strategy that separates speech from environmental sound and distinguishes multiple speakers, providing explicit perceptual reasoning for training. Building on this, we propose HyPeR, a two-stage Hybrid Perception-Reasoning framework. In Stage I, we finetune the model on PAQA to perceive acoustic attributes in complex audio. In Stage II, we leverage GRPO to refine the model's internal deliberation. We also introduce PAUSE tokens to facilitate latent computation during acoustically ambiguous phases and design perceptual consistency reward to align reasoning rationales with raw audio. Experiments across benchmarks demonstrate that HyPeR achieves absolute improvements over the base model, with performance comparable to large-scale models, stressing the effectiveness of hybrid perception-grounded reasoning for robust and multi-speaker audio understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding
Wang, Jieyi
Niu, Yazhe
Xu, Dexuan
Wei, Zhongyu
Sound
Multimedia
Recent Large Audio Language Models have demonstrated impressive capabilities in audio understanding. However, they often suffer from perceptual errors, while reliable audio reasoning is unattainable without first grounding the model's perception in structured auditory scenes. Inspired by Auditory Scene Analysis, we first introduce a Perception-Aware Question Answering (PAQA) dataset. PAQA implements a hierarchical decoupling strategy that separates speech from environmental sound and distinguishes multiple speakers, providing explicit perceptual reasoning for training. Building on this, we propose HyPeR, a two-stage Hybrid Perception-Reasoning framework. In Stage I, we finetune the model on PAQA to perceive acoustic attributes in complex audio. In Stage II, we leverage GRPO to refine the model's internal deliberation. We also introduce PAUSE tokens to facilitate latent computation during acoustically ambiguous phases and design perceptual consistency reward to align reasoning rationales with raw audio. Experiments across benchmarks demonstrate that HyPeR achieves absolute improvements over the base model, with performance comparable to large-scale models, stressing the effectiveness of hybrid perception-grounded reasoning for robust and multi-speaker audio understanding.
title Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding
topic Sound
Multimedia
url https://arxiv.org/abs/2604.14806