Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yifu, Ji, Shengpeng, Liu, Zhengqing, Chen, Qian, Wang, Wen, Wang, Ziqing, Li, Yangzhuo, Liang, Tianle, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
von: Chen, Yifu, et al.
Veröffentlicht: (2026)
von: Chen, Yifu, et al.
Veröffentlicht: (2026)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026)
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026)
VoxMind: An End-to-End Agentic Spoken Dialogue System
von: Liang, Tianle, et al.
Veröffentlicht: (2026)
von: Liang, Tianle, et al.
Veröffentlicht: (2026)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
von: Lu, Jingyu, et al.
Veröffentlicht: (2026)
von: Lu, Jingyu, et al.
Veröffentlicht: (2026)
WavChat: A Survey of Spoken Dialogue Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems
von: Elmers, Mikey, et al.
Veröffentlicht: (2025)
von: Elmers, Mikey, et al.
Veröffentlicht: (2025)
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025)
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
von: Li, Guojian, et al.
Veröffentlicht: (2025)
von: Li, Guojian, et al.
Veröffentlicht: (2025)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
von: Yan, Ruiqi, et al.
Veröffentlicht: (2025)
von: Yan, Ruiqi, et al.
Veröffentlicht: (2025)
JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems
von: Yang, Guangzhao, et al.
Veröffentlicht: (2026)
von: Yang, Guangzhao, et al.
Veröffentlicht: (2026)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
Aligning Spoken Dialogue Models from User Interactions
von: Wu, Anne, et al.
Veröffentlicht: (2025)
von: Wu, Anne, et al.
Veröffentlicht: (2025)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
Are LLMs Robust for Spoken Dialogues?
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
PRMB: Benchmarking Reward Models in Long-Horizon CBT-based Counseling Dialogue
von: Zhou, Yougen, et al.
Veröffentlicht: (2026)
von: Zhou, Yougen, et al.
Veröffentlicht: (2026)
TiCo: Time-Controllable Spoken Dialogue Model
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2026)
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2026)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
von: Gosai, Advait, et al.
Veröffentlicht: (2025)
von: Gosai, Advait, et al.
Veröffentlicht: (2025)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Applying General Turn-taking Models to Conversational Human-Robot Interaction
von: Skantze, Gabriel, et al.
Veröffentlicht: (2025)
von: Skantze, Gabriel, et al.
Veröffentlicht: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
von: Ivry, Amir, et al.
Veröffentlicht: (2026)
von: Ivry, Amir, et al.
Veröffentlicht: (2026)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
von: Qian, Livia, et al.
Veröffentlicht: (2024)
von: Qian, Livia, et al.
Veröffentlicht: (2024)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
von: Wan, Yanming, et al.
Veröffentlicht: (2025)
von: Wan, Yanming, et al.
Veröffentlicht: (2025)
Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025) -
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
von: Chen, Yifu, et al.
Veröffentlicht: (2026) -
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026) -
VoxMind: An End-to-End Agentic Spoken Dialogue System
von: Liang, Tianle, et al.
Veröffentlicht: (2026) -
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)