Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
Fuente:
arXiv
Saved in:
| Main Authors: | Arora, Siddhant, Khan, Haidar, Sun, Kai, Dong, Xin Luna, Choudhary, Sajal, Moon, Seungwhan, Zhang, Xinyuan, Sagar, Adithya, Appini, Surya Teja, Patnaik, Kaushik, Sharma, Sanat, Watanabe, Shinji, Kumar, Anuj, Aly, Ahmed, Liu, Yue, Metze, Florian, Lin, Zhaojiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
by: Lin, Zhaojiang, et al.
Published: (2025)
by: Lin, Zhaojiang, et al.
Published: (2025)
Semi-Autoregressive Streaming ASR With Label Context
by: Arora, Siddhant, et al.
Published: (2023)
by: Arora, Siddhant, et al.
Published: (2023)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
by: Udupa, Sathvik, et al.
Published: (2025)
by: Udupa, Sathvik, et al.
Published: (2025)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
by: Tsunoo, Emiru, et al.
Published: (2024)
by: Tsunoo, Emiru, et al.
Published: (2024)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
by: Arora, Siddhant, et al.
Published: (2026)
by: Arora, Siddhant, et al.
Published: (2026)
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
by: Li, Zekun, et al.
Published: (2024)
by: Li, Zekun, et al.
Published: (2024)
Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning
by: Chen, Jingxiang, et al.
Published: (2026)
by: Chen, Jingxiang, et al.
Published: (2026)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
by: Tsunoo, Emiru, et al.
Published: (2025)
by: Tsunoo, Emiru, et al.
Published: (2025)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
by: Futami, Hayato, et al.
Published: (2024)
by: Futami, Hayato, et al.
Published: (2024)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
by: Kim, Jeonghwan, et al.
Published: (2026)
by: Kim, Jeonghwan, et al.
Published: (2026)
Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systems
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
by: Le, Trang, et al.
Published: (2024)
by: Le, Trang, et al.
Published: (2024)
Instant Gaussian Stream: Fast and Generalizable Streaming of Dynamic Scene Reconstruction via Gaussian Splatting
by: Yan, Jinbo, et al.
Published: (2025)
by: Yan, Jinbo, et al.
Published: (2025)
Low-Latency Grid Intelligence with Self-Governing Stream and Calibration Agents
by: Parthasarathy, Adithya, et al.
Published: (2026)
by: Parthasarathy, Adithya, et al.
Published: (2026)
suryatejabatchu08/Parkinsonian-Gait-Assessment: PGSI: Pose-Derived Parkinsonian Gait Severity Index — Initial Release
by: Surya Teja
Published: (2026)
by: Surya Teja
Published: (2026)
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
by: Lu, Hui, et al.
Published: (2026)
by: Lu, Hui, et al.
Published: (2026)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
by: Arora, Siddhant, et al.
Published: (2023)
by: Arora, Siddhant, et al.
Published: (2023)
AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
by: Ivry, Amir, et al.
Published: (2026)
by: Ivry, Amir, et al.
Published: (2026)
On The Landscape of Spoken Language Models: A Comprehensive Survey
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
AraSpot: Arabic Spoken Command Spotting
by: Salhab, Mahmoud, et al.
Published: (2023)
by: Salhab, Mahmoud, et al.
Published: (2023)
Two-Stream Instability and Bernstein-Greene-Kruskal Mode Formation in Coulomb One Component Plasma
by: Mir, Ajaz, et al.
Published: (2025)
by: Mir, Ajaz, et al.
Published: (2025)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
by: Shakeel, Muhammad, et al.
Published: (2024)
by: Shakeel, Muhammad, et al.
Published: (2024)
Identity Control Plane: The Unifying Layer for Zero Trust Infrastructure
by: Avirneni, Surya Teja
Published: (2025)
by: Avirneni, Surya Teja
Published: (2025)
Establishing Workload Identity for Zero Trust CI/CD: From Secrets to SPIFFE-Based Authentication
by: Avirneni, Surya Teja
Published: (2025)
by: Avirneni, Surya Teja
Published: (2025)
Intent-Aware Authorization for Zero Trust CI/CD
by: Avirneni, Surya Teja
Published: (2025)
by: Avirneni, Surya Teja
Published: (2025)
Decoupling Identity from Access: Credential Broker Patterns for Secure CI/CD
by: Avirneni, Surya Teja
Published: (2025)
by: Avirneni, Surya Teja
Published: (2025)
LL-GABR: Energy Efficient Live Video Streaming Using Reinforcement Learning
by: Raman, Adithya, et al.
Published: (2024)
by: Raman, Adithya, et al.
Published: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
by: Arora, Siddhant, et al.
Published: (2024)
by: Arora, Siddhant, et al.
Published: (2024)
On-device Streaming Discrete Speech Units
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
Unit Interval Selection in Random Order Streams
by: Alexandru, Cezar-Mihail, et al.
Published: (2026)
by: Alexandru, Cezar-Mihail, et al.
Published: (2026)
Are LLMs Robust for Spoken Dialogues?
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
CoSMoEs: Compact Sparse Mixture of Experts
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
by: Nakata, Wataru, et al.
Published: (2024)
by: Nakata, Wataru, et al.
Published: (2024)
Comparative Analysis of Data Warehousing Solutions: AWS Redshift vs. Snowflake vs. Google BigQuery
by: Naga Surya Teja Thallam
Published: (2020)
by: Naga Surya Teja Thallam
Published: (2020)
Similar Items
-
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025) -
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025) -
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
by: Lin, Zhaojiang, et al.
Published: (2025) -
Semi-Autoregressive Streaming ASR With Label Context
by: Arora, Siddhant, et al.
Published: (2023) -
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
by: Udupa, Sathvik, et al.
Published: (2025)