Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Siyuan, Xu, Jiahui, Jiang, Feng, Wang, Kuang, Zhao, Zefeng, Huang, Chu-Ren, Gu, Jinghang, Yin, Changqing, Li, Haizhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unsupervised Mutual Learning of Discourse Parsing and Topic Segmentation in Dialogue
di: Xu, Jiahui, et al.
Pubblicazione: (2024)
di: Xu, Jiahui, et al.
Pubblicazione: (2024)
Lightweight Multimodal Edge Computing for Low‐Latency Dialogue Simulation in Spoken English Practice
di: Ying Huang
Pubblicazione: (2025)
di: Ying Huang
Pubblicazione: (2025)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical Study
di: Fan, Yaxin, et al.
Pubblicazione: (2023)
di: Fan, Yaxin, et al.
Pubblicazione: (2023)
CHisAgent: A Multi-Agent Framework for Event Taxonomy Construction in Ancient Chinese Cultural Systems
di: Tang, Xuemei, et al.
Pubblicazione: (2026)
di: Tang, Xuemei, et al.
Pubblicazione: (2026)
Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Bridging Research and Readers: A Multi-Modal Automated Academic Papers Interpretation System
di: Jiang, Feng, et al.
Pubblicazione: (2024)
di: Jiang, Feng, et al.
Pubblicazione: (2024)
Joint Information Extraction Across Classical and Modern Chinese with Tea-MOELoRA
di: Tang, Xuemei, et al.
Pubblicazione: (2025)
di: Tang, Xuemei, et al.
Pubblicazione: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Towards a Japanese Full-duplex Spoken Dialogue System
di: Ohashi, Atsumoto, et al.
Pubblicazione: (2025)
di: Ohashi, Atsumoto, et al.
Pubblicazione: (2025)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
Human Latency Conversational Turns for Spoken Avatar Systems
di: Jacoby, Derek, et al.
Pubblicazione: (2024)
di: Jacoby, Derek, et al.
Pubblicazione: (2024)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
di: Ao, Junyi, et al.
Pubblicazione: (2024)
di: Ao, Junyi, et al.
Pubblicazione: (2024)
Towards Low-Latency Event Stream-based Visual Object Tracking: A Slow-Fast Approach
di: Wang, Shiao, et al.
Pubblicazione: (2025)
di: Wang, Shiao, et al.
Pubblicazione: (2025)
Data Augmentation Integrating Dialogue Flow and Style to Adapt Spoken Dialogue Systems to Low-Resource User Groups
di: Qi, Zhiyang, et al.
Pubblicazione: (2024)
di: Qi, Zhiyang, et al.
Pubblicazione: (2024)
CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation
di: Ke, Rui, et al.
Pubblicazione: (2025)
di: Ke, Rui, et al.
Pubblicazione: (2025)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
di: Zhang, Hao, et al.
Pubblicazione: (2025)
di: Zhang, Hao, et al.
Pubblicazione: (2025)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
di: Peng, Yizhou, et al.
Pubblicazione: (2026)
di: Peng, Yizhou, et al.
Pubblicazione: (2026)
WavChat: A Survey of Spoken Dialogue Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
MOSS-TTSD: Text to Spoken Dialogue Generation
di: Zhang, Yuqian, et al.
Pubblicazione: (2026)
di: Zhang, Yuqian, et al.
Pubblicazione: (2026)
Low-Latency Neural Stereo Streaming
di: Hou, Qiqi, et al.
Pubblicazione: (2024)
di: Hou, Qiqi, et al.
Pubblicazione: (2024)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking
di: Hu, Guyue, et al.
Pubblicazione: (2026)
di: Hu, Guyue, et al.
Pubblicazione: (2026)
Are LLMs Robust for Spoken Dialogues?
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2024)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
Investigating Low-Cost LLM Annotation for~Spoken Dialogue Understanding Datasets
di: Druart, Lucas, et al.
Pubblicazione: (2024)
di: Druart, Lucas, et al.
Pubblicazione: (2024)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
di: Liu, Jingwen, et al.
Pubblicazione: (2025)
di: Liu, Jingwen, et al.
Pubblicazione: (2025)
UNO-DST: Leveraging Unlabelled Data in Zero-Shot Dialogue State Tracking
di: Li, Chuang, et al.
Pubblicazione: (2023)
di: Li, Chuang, et al.
Pubblicazione: (2023)
Dialogue and Discourse
Pubblicazione: (2017)
Pubblicazione: (2017)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
di: Lin, Guan-Ting, et al.
Pubblicazione: (2023)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2023)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems
di: Inoue, Koji, et al.
Pubblicazione: (2024)
di: Inoue, Koji, et al.
Pubblicazione: (2024)
VoxMind: An End-to-End Agentic Spoken Dialogue System
di: Liang, Tianle, et al.
Pubblicazione: (2026)
di: Liang, Tianle, et al.
Pubblicazione: (2026)
SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
di: Lee, Jonggeun, et al.
Pubblicazione: (2026)
di: Lee, Jonggeun, et al.
Pubblicazione: (2026)
Paralinguistic Emotion-Aware Validation Timing Detection in Japanese Empathetic Spoken Dialogue
di: Pang, Zi Haur, et al.
Pubblicazione: (2026)
di: Pang, Zi Haur, et al.
Pubblicazione: (2026)
Adapting Text-based Dialogue State Tracker for Spoken Dialogues
di: Yoon, Jaeseok, et al.
Pubblicazione: (2023)
di: Yoon, Jaeseok, et al.
Pubblicazione: (2023)
SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
di: Han, Jionghao, et al.
Pubblicazione: (2025)
di: Han, Jionghao, et al.
Pubblicazione: (2025)
Staircase Streaming for Low-Latency Multi-Agent Inference
di: Wang, Junlin, et al.
Pubblicazione: (2025)
di: Wang, Junlin, et al.
Pubblicazione: (2025)
Low-Latency Scalable Streaming for Event-Based Vision
di: Hamara, Andrew, et al.
Pubblicazione: (2024)
di: Hamara, Andrew, et al.
Pubblicazione: (2024)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
di: Vendrame, Katia, et al.
Pubblicazione: (2025)
di: Vendrame, Katia, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unsupervised Mutual Learning of Discourse Parsing and Topic Segmentation in Dialogue
di: Xu, Jiahui, et al.
Pubblicazione: (2024) -
Lightweight Multimodal Edge Computing for Low‐Latency Dialogue Simulation in Spoken English Practice
di: Ying Huang
Pubblicazione: (2025) -
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024) -
Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical Study
di: Fan, Yaxin, et al.
Pubblicazione: (2023) -
CHisAgent: A Multi-Agent Framework for Event Taxonomy Construction in Ancient Chinese Cultural Systems
di: Tang, Xuemei, et al.
Pubblicazione: (2026)