Saved in:
Bibliographic Details
Main Authors: Song, Zhiyuan, Zhao, Weici, Xiao, Yang, Yu, Suhao, Zhu, Cheng, Gu, Jiatao
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.27190
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914605072121856
author Song, Zhiyuan
Zhao, Weici
Xiao, Yang
Yu, Suhao
Zhu, Cheng
Gu, Jiatao
author_facet Song, Zhiyuan
Zhao, Weici
Xiao, Yang
Yu, Suhao
Zhu, Cheng
Gu, Jiatao
contents Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsiveness are tightly coupled: delaying reasoning until the speech endpoint can improve answer quality but moves deliberation into user-visible response delay, while answering too early risks committing before decisive evidence arrives. We introduce a learnable wait-think-answer control formulation for LALMs. Motivated by the incremental nature of human conversation, the controller decides under partial audio evidence when to wait, when to externalize a compact reasoning update, and when to answer. Using Qwen2.5-Omni-7B as the base model, we construct aligned wait-think-answer traces from spoken reasoning data, train the controller with supervised fine-tuning (SFT), and then apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO). The reward combines answer correctness, action validity, update timing, latency synchronization, reasoning quality, and chain consistency, optimizing the complete wait-think-answer trajectory and not the final answer alone. On a six-task synthetic spoken reasoning question answering (SRQA) benchmark, the six-reward DAPO controller improves the row-weighted accuracy from 67.6% to 70.3% while reducing post-endpoint final-think length by 14% under the same Qwen deployment harness. On a 186-item human-recorded Real Audio Bench, a transfer check beyond text-to-speech (TTS)-rendered speech, the controller family remains functional: SFT achieves the strongest accuracy, while the six-reward DAPO controller is the only learned variant whose final-think length falls below the base. These results suggest that a streaming model should learn when to make intermediate reasoning explicit during the audio stream.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27190
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning When to Think While Listening in Large Audio-Language Models
Song, Zhiyuan
Zhao, Weici
Xiao, Yang
Yu, Suhao
Zhu, Cheng
Gu, Jiatao
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsiveness are tightly coupled: delaying reasoning until the speech endpoint can improve answer quality but moves deliberation into user-visible response delay, while answering too early risks committing before decisive evidence arrives. We introduce a learnable wait-think-answer control formulation for LALMs. Motivated by the incremental nature of human conversation, the controller decides under partial audio evidence when to wait, when to externalize a compact reasoning update, and when to answer. Using Qwen2.5-Omni-7B as the base model, we construct aligned wait-think-answer traces from spoken reasoning data, train the controller with supervised fine-tuning (SFT), and then apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO). The reward combines answer correctness, action validity, update timing, latency synchronization, reasoning quality, and chain consistency, optimizing the complete wait-think-answer trajectory and not the final answer alone. On a six-task synthetic spoken reasoning question answering (SRQA) benchmark, the six-reward DAPO controller improves the row-weighted accuracy from 67.6% to 70.3% while reducing post-endpoint final-think length by 14% under the same Qwen deployment harness. On a 186-item human-recorded Real Audio Bench, a transfer check beyond text-to-speech (TTS)-rendered speech, the controller family remains functional: SFT achieves the strongest accuracy, while the six-reward DAPO controller is the only learned variant whose final-think length falls below the base. These results suggest that a streaming model should learn when to make intermediate reasoning explicit during the audio stream.
title Learning When to Think While Listening in Large Audio-Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2605.27190