RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jing-Han, Su, Bo-Hao, Wu, Ya-Tse, Lee, Chi-Chun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910018921562112
author Chen, Jing-Han
Su, Bo-Hao
Wu, Ya-Tse
Lee, Chi-Chun
author_facet Chen, Jing-Han
Su, Bo-Hao
Wu, Ya-Tse
Lee, Chi-Chun
contents With generative AI advancing, empathy in human-AI interaction is essential. While prior work focuses on emotional reflection, emotional exploration, key to deeper engagement, remains overlooked. Existing LLMs rely on text which captures limited emotion nuances. To address this, we propose RE-LLM, a speech-LLM integrating dimensional emotion embeddings and auxiliary learning. Experiments show statistically significant gains in empathy metrics across three datasets. RE-LLM relatively improves the Emotional Reaction score by 14.79% and 6.76% compared to text-only and speech-LLM baselines on ESD. Notably, it raises the Exploration score by 35.42% and 3.91% on IEMOCAP, 139.28% and 9.83% on ESD, and 60.95% and 22.64% on MSP-PODCAST. It also boosts unweighted accuracy by 5.4% on IEMOCAP, 2.3% on ESD, and 6.9% on MSP-PODCAST in speech emotion recognition. These results highlight the enriched emotional understanding and improved empathetic response generation of RE-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10716
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance
Chen, Jing-Han
Su, Bo-Hao
Wu, Ya-Tse
Lee, Chi-Chun
Audio and Speech Processing
Computation and Language
Sound
With generative AI advancing, empathy in human-AI interaction is essential. While prior work focuses on emotional reflection, emotional exploration, key to deeper engagement, remains overlooked. Existing LLMs rely on text which captures limited emotion nuances. To address this, we propose RE-LLM, a speech-LLM integrating dimensional emotion embeddings and auxiliary learning. Experiments show statistically significant gains in empathy metrics across three datasets. RE-LLM relatively improves the Emotional Reaction score by 14.79% and 6.76% compared to text-only and speech-LLM baselines on ESD. Notably, it raises the Exploration score by 35.42% and 3.91% on IEMOCAP, 139.28% and 9.83% on ESD, and 60.95% and 22.64% on MSP-PODCAST. It also boosts unweighted accuracy by 5.4% on IEMOCAP, 2.3% on ESD, and 6.9% on MSP-PODCAST in speech emotion recognition. These results highlight the enriched emotional understanding and improved empathetic response generation of RE-LLM.
title RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2602.10716