EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fei, Hao, Zhang, Han, Wang, Bin, Liao, Lizi, Liu, Qian, Cambria, Erik
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929394320146432
author Fei, Hao
Zhang, Han
Wang, Bin
Liao, Lizi
Liu, Qian
Cambria, Erik
author_facet Fei, Hao
Zhang, Han
Wang, Bin
Liao, Lizi
Liu, Qian
Cambria, Erik
contents This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating human-like empathy. The system paves the way for the next emotional intelligence, for which we open-source the code for public access.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15177
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
Fei, Hao
Zhang, Han
Wang, Bin
Liao, Lizi
Liu, Qian
Cambria, Erik
Multimedia
This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating human-like empathy. The system paves the way for the next emotional intelligence, for which we open-source the code for public access.
title EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
topic Multimedia
url https://arxiv.org/abs/2406.15177