Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arora, Siddhant, Khan, Haidar, Sun, Kai, Dong, Xin Luna, Choudhary, Sajal, Moon, Seungwhan, Zhang, Xinyuan, Sagar, Adithya, Appini, Surya Teja, Patnaik, Kaushik, Sharma, Sanat, Watanabe, Shinji, Kumar, Anuj, Aly, Ahmed, Liu, Yue, Metze, Florian, Lin, Zhaojiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911189115600896
author Arora, Siddhant
Khan, Haidar
Sun, Kai
Dong, Xin Luna
Choudhary, Sajal
Moon, Seungwhan
Zhang, Xinyuan
Sagar, Adithya
Appini, Surya Teja
Patnaik, Kaushik
Sharma, Sanat
Watanabe, Shinji
Kumar, Anuj
Aly, Ahmed
Liu, Yue
Metze, Florian
Lin, Zhaojiang
author_facet Arora, Siddhant
Khan, Haidar
Sun, Kai
Dong, Xin Luna
Choudhary, Sajal
Moon, Seungwhan
Zhang, Xinyuan
Sagar, Adithya
Appini, Surya Teja
Patnaik, Kaushik
Sharma, Sanat
Watanabe, Shinji
Kumar, Anuj
Aly, Ahmed
Liu, Yue
Metze, Florian
Lin, Zhaojiang
contents End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with significantly lower latency. However, these systems remain prone to hallucinations due to limited factual grounding. While text-based dialogue systems address this challenge by integrating tools such as web search and knowledge graph APIs, we introduce the first approach to extend tool use directly into speech-in speech-out systems. A key challenge is that tool integration substantially increases response latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Streaming RAG), a novel framework that reduces user-perceived latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. Specifically, we develop a post-training pipeline that teaches the model when to issue tool calls during ongoing speech and how to generate spoken summaries that fuse audio queries with retrieved text results, thereby improving both accuracy and responsiveness. To evaluate our approach, we construct AudioCRAG, a benchmark created by converting queries from the publicly available CRAG dataset into speech form. Experimental results demonstrate that our streaming RAG approach increases QA accuracy by up to 200% relative (from 11.1% to 34.2% absolute) and further enhances user experience by reducing tool use latency by 20%. Importantly, our streaming RAG approach is modality-agnostic and can be applied equally to typed input, paving the way for more agentic, real-time AI assistants.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02044
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
Arora, Siddhant
Khan, Haidar
Sun, Kai
Dong, Xin Luna
Choudhary, Sajal
Moon, Seungwhan
Zhang, Xinyuan
Sagar, Adithya
Appini, Surya Teja
Patnaik, Kaushik
Sharma, Sanat
Watanabe, Shinji
Kumar, Anuj
Aly, Ahmed
Liu, Yue
Metze, Florian
Lin, Zhaojiang
Computation and Language
Sound
Audio and Speech Processing
End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with significantly lower latency. However, these systems remain prone to hallucinations due to limited factual grounding. While text-based dialogue systems address this challenge by integrating tools such as web search and knowledge graph APIs, we introduce the first approach to extend tool use directly into speech-in speech-out systems. A key challenge is that tool integration substantially increases response latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Streaming RAG), a novel framework that reduces user-perceived latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. Specifically, we develop a post-training pipeline that teaches the model when to issue tool calls during ongoing speech and how to generate spoken summaries that fuse audio queries with retrieved text results, thereby improving both accuracy and responsiveness. To evaluate our approach, we construct AudioCRAG, a benchmark created by converting queries from the publicly available CRAG dataset into speech form. Experimental results demonstrate that our streaming RAG approach increases QA accuracy by up to 200% relative (from 11.1% to 34.2% absolute) and further enhances user experience by reducing tool use latency by 20%. Importantly, our streaming RAG approach is modality-agnostic and can be applied equally to typed input, paving the way for more agentic, real-time AI assistants.
title Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.02044