E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Ronghao, Shen, Shuai, Hu, Weipeng, He, Qiaolin, Xiong, Aolin, Huang, Li, Hu, Haifeng, Tan, Yap-peng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913996239536128
author Lin, Ronghao
Shen, Shuai
Hu, Weipeng
He, Qiaolin
Xiong, Aolin
Huang, Li
Hu, Haifeng
Tan, Yap-peng
author_facet Lin, Ronghao
Shen, Shuai
Hu, Weipeng
He, Qiaolin
Xiong, Aolin
Huang, Li
Hu, Haifeng
Tan, Yap-peng
contents Multimodal Empathetic Response Generation (MERG) is crucial for building emotionally intelligent human-computer interactions. Although large language models (LLMs) have improved text-based ERG, challenges remain in handling multimodal emotional content and maintaining identity consistency. Thus, we propose E3RG, an Explicit Emotion-driven Empathetic Response Generation System based on multimodal LLMs which decomposes MERG task into three parts: multimodal empathy understanding, empathy memory retrieval, and multimodal response generation. By integrating advanced expressive speech and video generative models, E3RG delivers natural, emotionally rich, and identity-consistent responses without extra training. Experiments validate the superiority of our system on both zero-shot and few-shot settings, securing Top-1 position in the Avatar-based Multimodal Empathy Challenge on ACM MM 25. Our code is available at https://github.com/RH-Lin/E3RG.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Lin, Ronghao
Shen, Shuai
Hu, Weipeng
He, Qiaolin
Xiong, Aolin
Huang, Li
Hu, Haifeng
Tan, Yap-peng
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Human-Computer Interaction
Multimedia
Multimodal Empathetic Response Generation (MERG) is crucial for building emotionally intelligent human-computer interactions. Although large language models (LLMs) have improved text-based ERG, challenges remain in handling multimodal emotional content and maintaining identity consistency. Thus, we propose E3RG, an Explicit Emotion-driven Empathetic Response Generation System based on multimodal LLMs which decomposes MERG task into three parts: multimodal empathy understanding, empathy memory retrieval, and multimodal response generation. By integrating advanced expressive speech and video generative models, E3RG delivers natural, emotionally rich, and identity-consistent responses without extra training. Experiments validate the superiority of our system on both zero-shot and few-shot settings, securing Top-1 position in the Avatar-based Multimodal Empathy Challenge on ACM MM 25. Our code is available at https://github.com/RH-Lin/E3RG.
title E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Human-Computer Interaction
Multimedia
url https://arxiv.org/abs/2508.12854