Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Riera, Pablo, Brusco, Pablo, Kuo, Cristina, Sancinetti, Marcelo, Branavan, S. R. K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
von: Li, Guojian, et al.
Veröffentlicht: (2025)
von: Li, Guojian, et al.
Veröffentlicht: (2025)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
von: Yao, Yiqun, et al.
Veröffentlicht: (2025)
von: Yao, Yiqun, et al.
Veröffentlicht: (2025)
JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems
von: Yang, Guangzhao, et al.
Veröffentlicht: (2026)
von: Yang, Guangzhao, et al.
Veröffentlicht: (2026)
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
von: Modi, Smit Nautambhai, et al.
Veröffentlicht: (2026)
von: Modi, Smit Nautambhai, et al.
Veröffentlicht: (2026)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
MOSS-TTSD: Text to Spoken Dialogue Generation
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
Efficient Training for Cross-lingual Speech Language Models
von: Zhou, Yan, et al.
Veröffentlicht: (2026)
von: Zhou, Yan, et al.
Veröffentlicht: (2026)
Cross-Attention is Half Explanation in Speech-to-Text Models
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
von: Liu, Ruohan, et al.
Veröffentlicht: (2026)
von: Liu, Ruohan, et al.
Veröffentlicht: (2026)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
von: Rosero, Karen, et al.
Veröffentlicht: (2025)
von: Rosero, Karen, et al.
Veröffentlicht: (2025)
Raon-Speech Technical Report
von: Kim, Beomsoo, et al.
Veröffentlicht: (2026)
von: Kim, Beomsoo, et al.
Veröffentlicht: (2026)
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
von: Gedeon, Máté, et al.
Veröffentlicht: (2025)
von: Gedeon, Máté, et al.
Veröffentlicht: (2025)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout Fidelity
von: Hammer, Hugo L., et al.
Veröffentlicht: (2026)
von: Hammer, Hugo L., et al.
Veröffentlicht: (2026)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
von: Ginjala, Srishti, et al.
Veröffentlicht: (2026)
von: Ginjala, Srishti, et al.
Veröffentlicht: (2026)
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
von: Božík, Erik, et al.
Veröffentlicht: (2025)
von: Božík, Erik, et al.
Veröffentlicht: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild
von: Durmus, Berkin, et al.
Veröffentlicht: (2026)
von: Durmus, Berkin, et al.
Veröffentlicht: (2026)
MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning
von: Zheng, Haolong, et al.
Veröffentlicht: (2026)
von: Zheng, Haolong, et al.
Veröffentlicht: (2026)
Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care
von: Lyberatos, Vassilis, et al.
Veröffentlicht: (2026)
von: Lyberatos, Vassilis, et al.
Veröffentlicht: (2026)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
von: Gao, Yan, et al.
Veröffentlicht: (2025)
von: Gao, Yan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025) -
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024) -
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
von: Li, Guojian, et al.
Veröffentlicht: (2025) -
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026) -
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
von: Yao, Yiqun, et al.
Veröffentlicht: (2025)