Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shih, Yi-Jen, Raj, Desh, Wu, Chunyang, Zhou, Wei, Bong, SK, Gaur, Yashesh, Mahadeokar, Jay, Kalinli, Ozlem, Seltzer, Mike
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2510.07497
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916997862785024
author Shih, Yi-Jen
Raj, Desh
Wu, Chunyang
Zhou, Wei
Bong, SK
Gaur, Yashesh
Mahadeokar, Jay
Kalinli, Ozlem
Seltzer, Mike
author_facet Shih, Yi-Jen
Raj, Desh
Wu, Chunyang
Zhou, Wei
Bong, SK
Gaur, Yashesh
Mahadeokar, Jay
Kalinli, Ozlem
Seltzer, Mike
contents Recent advances in speech large language models (speech LLMs) have enabled seamless spoken interactions, but these systems still struggle with complex reasoning tasks. Previously, chain-of-thought (CoT) prompting or fine-tuning has been to shown to significantly improve the reasoning abilities of text-based LLMs. In this work, we investigate the effect of CoT fine-tuning for multi-stream speech LLMs, demonstrating that reasoning in text space improves the accuracy of speech LLMs by 2.4x, on average, over a suite of spoken reasoning tasks. Beyond accuracy, the latency of the spoken response is a crucial factor for interacting with voice-based agents. Inspired by the human behavior of "thinking while listening," we propose methods to reduce the additional latency from reasoning by allowing the model to start reasoning before the user query has ended. To achieve this, we introduce an entropy-based metric, "question completeness," which acts as an indicator to guide the model on the optimal time to start reasoning. This method provides greater control over the accuracy-latency trade-off compared with heuristic-based approaches and, under equivalent latency conditions, yields a 4% accuracy gain on ARC-Easy. Finally, we use Direct Preference Optimization (DPO) on preference data created using rejection sampling to push the accuracy-latency pareto frontier further, resulting in a 70% reduction in latency without loss in accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Speech LLMs Think while Listening?
Shih, Yi-Jen
Raj, Desh
Wu, Chunyang
Zhou, Wei
Bong, SK
Gaur, Yashesh
Mahadeokar, Jay
Kalinli, Ozlem
Seltzer, Mike
Computation and Language
Artificial Intelligence
Audio and Speech Processing
Recent advances in speech large language models (speech LLMs) have enabled seamless spoken interactions, but these systems still struggle with complex reasoning tasks. Previously, chain-of-thought (CoT) prompting or fine-tuning has been to shown to significantly improve the reasoning abilities of text-based LLMs. In this work, we investigate the effect of CoT fine-tuning for multi-stream speech LLMs, demonstrating that reasoning in text space improves the accuracy of speech LLMs by 2.4x, on average, over a suite of spoken reasoning tasks. Beyond accuracy, the latency of the spoken response is a crucial factor for interacting with voice-based agents. Inspired by the human behavior of "thinking while listening," we propose methods to reduce the additional latency from reasoning by allowing the model to start reasoning before the user query has ended. To achieve this, we introduce an entropy-based metric, "question completeness," which acts as an indicator to guide the model on the optimal time to start reasoning. This method provides greater control over the accuracy-latency trade-off compared with heuristic-based approaches and, under equivalent latency conditions, yields a 4% accuracy gain on ARC-Easy. Finally, we use Direct Preference Optimization (DPO) on preference data created using rejection sampling to push the accuracy-latency pareto frontier further, resulting in a 70% reduction in latency without loss in accuracy.
title Can Speech LLMs Think while Listening?
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.07497