OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
Fuente:
arXiv
Saved in:
| Main Authors: | Geng, Xuelong, Shao, Qijie, Xue, Hongfei, Wang, Shuiyuan, Xie, Hanke, Guo, Zhao, Zhao, Yi, Li, Guojian, Tian, Wenjie, Wang, Chengyou, Zhao, Zhixian, Xia, Kangxiang, Zhang, Ziyu, Lin, Zhennan, Zuo, Tianlun, Shao, Mingchen, Cao, Yuang, Ma, Guobin, Li, Longhao, Dai, Yuhang, Gao, Dehui, Guo, Dake, Xie, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
by: Geng, Xuelong, et al.
Published: (2025)
by: Geng, Xuelong, et al.
Published: (2025)
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge
by: Wang, Chengyou, et al.
Published: (2026)
by: Wang, Chengyou, et al.
Published: (2026)
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
by: Li, Guojian, et al.
Published: (2025)
by: Li, Guojian, et al.
Published: (2025)
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
by: Li, Guojian, et al.
Published: (2025)
by: Li, Guojian, et al.
Published: (2025)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
by: Liao, Yujie, et al.
Published: (2026)
by: Liao, Yujie, et al.
Published: (2026)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
by: Guo, Zhao, et al.
Published: (2025)
by: Guo, Zhao, et al.
Published: (2025)
FormalASR: End-to-End Spoken Chinese to Formal Text
by: Ning, Wanyi, et al.
Published: (2026)
by: Ning, Wanyi, et al.
Published: (2026)
VoxMind: An End-to-End Agentic Spoken Dialogue System
by: Liang, Tianle, et al.
Published: (2026)
by: Liang, Tianle, et al.
Published: (2026)
End-to-End Spoken Grammatical Error Correction
by: Qian, Mengjie, et al.
Published: (2025)
by: Qian, Mengjie, et al.
Published: (2025)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
by: Xia, Kangxiang, et al.
Published: (2026)
by: Xia, Kangxiang, et al.
Published: (2026)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
by: Yan, Ruiqi, et al.
Published: (2025)
by: Yan, Ruiqi, et al.
Published: (2025)
Privacy-Preserving End-to-End Spoken Language Understanding
by: Wang, Yinggui, et al.
Published: (2024)
by: Wang, Yinggui, et al.
Published: (2024)
Retrieval Augmented End-to-End Spoken Dialog Models
by: Wang, Mingqiu, et al.
Published: (2024)
by: Wang, Mingqiu, et al.
Published: (2024)
Towards End-to-End Spoken Grammatical Error Correction
by: Bannò, Stefano, et al.
Published: (2023)
by: Bannò, Stefano, et al.
Published: (2023)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
by: Li, Yangzhuo, et al.
Published: (2026)
by: Li, Yangzhuo, et al.
Published: (2026)
GSQA: An End-to-End Model for Generative Spoken Question Answering
by: Shih, Min-Han, et al.
Published: (2023)
by: Shih, Min-Han, et al.
Published: (2023)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
by: Tian, Wenjie, et al.
Published: (2026)
by: Tian, Wenjie, et al.
Published: (2026)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation
by: Sildam, Tiia, et al.
Published: (2024)
by: Sildam, Tiia, et al.
Published: (2024)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
by: Qian, Mengjie, et al.
Published: (2025)
by: Qian, Mengjie, et al.
Published: (2025)
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
by: Wang, Shuiyuan, et al.
Published: (2026)
by: Wang, Shuiyuan, et al.
Published: (2026)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
by: Zhao, Zhixian, et al.
Published: (2025)
by: Zhao, Zhixian, et al.
Published: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
by: Tian, Wenjie, et al.
Published: (2026)
by: Tian, Wenjie, et al.
Published: (2026)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
by: Le, Trang, et al.
Published: (2024)
by: Le, Trang, et al.
Published: (2024)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
by: Xie, Jingran, et al.
Published: (2025)
by: Xie, Jingran, et al.
Published: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
by: Hsiao, Chi-Yuan, et al.
Published: (2025)
by: Hsiao, Chi-Yuan, et al.
Published: (2025)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
by: Cui, Wenqian, et al.
Published: (2025)
by: Cui, Wenqian, et al.
Published: (2025)
Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
by: Li, Guojian, et al.
Published: (2026)
by: Li, Guojian, et al.
Published: (2026)
FreezeEmpath: Efficient Training for Empathetic Spoken Chatbots with Frozen LLMs
by: Hong, Yun, et al.
Published: (2026)
by: Hong, Yun, et al.
Published: (2026)
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
by: Hu, Jingbin, et al.
Published: (2026)
by: Hu, Jingbin, et al.
Published: (2026)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
by: Vendrame, Katia, et al.
Published: (2025)
by: Vendrame, Katia, et al.
Published: (2025)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
by: He, Jianfeng, et al.
Published: (2023)
by: He, Jianfeng, et al.
Published: (2023)
Paralinguistic Emotion-Aware Validation Timing Detection in Japanese Empathetic Spoken Dialogue
by: Pang, Zi Haur, et al.
Published: (2026)
by: Pang, Zi Haur, et al.
Published: (2026)
The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge
by: Guo, Dake, et al.
Published: (2024)
by: Guo, Dake, et al.
Published: (2024)
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
by: Ghazal, Nizar El, et al.
Published: (2025)
by: Ghazal, Nizar El, et al.
Published: (2025)
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
by: Dai, Yuhang, et al.
Published: (2025)
by: Dai, Yuhang, et al.
Published: (2025)
Similar Items
-
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
by: Geng, Xuelong, et al.
Published: (2025) -
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge
by: Wang, Chengyou, et al.
Published: (2026) -
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
by: Li, Guojian, et al.
Published: (2025) -
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
by: Li, Guojian, et al.
Published: (2025) -
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
by: Zhao, Zhixian, et al.
Published: (2026)