Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Haoyu, Zhang, Guangyan, Chen, Jiale, Li, Jingyu, Wang, Yuehai, Guo, Yiwen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915498839506944
author Wang, Haoyu
Zhang, Guangyan
Chen, Jiale
Li, Jingyu
Wang, Yuehai
Guo, Yiwen
author_facet Wang, Haoyu
Zhang, Guangyan
Chen, Jiale
Li, Jingyu
Wang, Yuehai
Guo, Yiwen
contents With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich emotional cues in user queries, where the same sentence may convey different meanings depending on the expression. Emotional understanding is thus essential for improving human-machine interaction. Most empathetic speech LLMs rely on massive datasets, demanding high computational cost. A key challenge is to build models that generate empathetic responses with limited data and without large-scale training. To this end, we propose Emotion Omni, a model that understands emotional content in user speech and generates empathetic responses. We further developed a data pipeline to construct a 200k emotional dialogue dataset supporting empathetic speech assistants. Experiments show that Emotion Omni achieves comparable instruction-following ability without large-scale pretraining, while surpassing existing models in speech quality (UTMOS:4.41) and empathy (Emotion GPT Score: 3.97). These results confirm its improvements in both speech fidelity and emotional expressiveness. Demos are available at https://w311411.github.io/omni_demo/.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18655
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
Wang, Haoyu
Zhang, Guangyan
Chen, Jiale
Li, Jingyu
Wang, Yuehai
Guo, Yiwen
Computation and Language
Sound
Audio and Speech Processing
I.2.7
With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich emotional cues in user queries, where the same sentence may convey different meanings depending on the expression. Emotional understanding is thus essential for improving human-machine interaction. Most empathetic speech LLMs rely on massive datasets, demanding high computational cost. A key challenge is to build models that generate empathetic responses with limited data and without large-scale training. To this end, we propose Emotion Omni, a model that understands emotional content in user speech and generates empathetic responses. We further developed a data pipeline to construct a 200k emotional dialogue dataset supporting empathetic speech assistants. Experiments show that Emotion Omni achieves comparable instruction-following ability without large-scale pretraining, while surpassing existing models in speech quality (UTMOS:4.41) and empathy (Emotion GPT Score: 3.97). These results confirm its improvements in both speech fidelity and emotional expressiveness. Demos are available at https://w311411.github.io/omni_demo/.
title Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
topic Computation and Language
Sound
Audio and Speech Processing
I.2.7
url https://arxiv.org/abs/2508.18655