Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Yuhang, Liu, Pei, Sun, Haoqin, Zhou, Jiaming, Cheng, Xuxin, Liu, Cao, Zeng, Ke, Cai, Xunliang, Qin, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026)
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
VoxMind: An End-to-End Agentic Spoken Dialogue System
von: Liang, Tianle, et al.
Veröffentlicht: (2026)
von: Liang, Tianle, et al.
Veröffentlicht: (2026)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025)
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025)
DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition
von: Liu, HongYu, et al.
Veröffentlicht: (2025)
von: Liu, HongYu, et al.
Veröffentlicht: (2025)
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
von: Chen, Tanyu, et al.
Veröffentlicht: (2026)
von: Chen, Tanyu, et al.
Veröffentlicht: (2026)
Retrieval Augmented End-to-End Spoken Dialog Models
von: Wang, Mingqiu, et al.
Veröffentlicht: (2024)
von: Wang, Mingqiu, et al.
Veröffentlicht: (2024)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
von: Guo, Yujie, et al.
Veröffentlicht: (2025)
von: Guo, Yujie, et al.
Veröffentlicht: (2025)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
von: Le, Trang, et al.
Veröffentlicht: (2024)
von: Le, Trang, et al.
Veröffentlicht: (2024)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2025)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Paralinguistic Emotion-Aware Validation Timing Detection in Japanese Empathetic Spoken Dialogue
von: Pang, Zi Haur, et al.
Veröffentlicht: (2026)
von: Pang, Zi Haur, et al.
Veröffentlicht: (2026)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Mean Opinion Score Prediction
von: Wang, Hui, et al.
Veröffentlicht: (2024)
von: Wang, Hui, et al.
Veröffentlicht: (2024)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
von: Han, Jionghao, et al.
Veröffentlicht: (2025)
von: Han, Jionghao, et al.
Veröffentlicht: (2025)
Pushing the Limits of End-to-End Diarization
von: Broughton, Samuel J., et al.
Veröffentlicht: (2025)
von: Broughton, Samuel J., et al.
Veröffentlicht: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
von: Cai, Danwei, et al.
Veröffentlicht: (2024)
von: Cai, Danwei, et al.
Veröffentlicht: (2024)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026) -
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
von: Geng, Xuelong, et al.
Veröffentlicht: (2025) -
VoxMind: An End-to-End Agentic Spoken Dialogue System
von: Liang, Tianle, et al.
Veröffentlicht: (2026) -
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025) -
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)