AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Se Jin, Kim, Yeonju, Rha, Hyeongseop, Godiva, Bella, Ro, Yong Man
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912166598148096
author Park, Se Jin
Kim, Yeonju
Rha, Hyeongseop
Godiva, Bella
Ro, Yong Man
author_facet Park, Se Jin
Kim, Yeonju
Rha, Hyeongseop
Godiva, Bella
Ro, Yong Man
contents In human communication, both verbal and non-verbal cues play a crucial role in conveying emotions, intentions, and meaning beyond words alone. These non-linguistic information, such as facial expressions, eye contact, voice tone, and pitch, are fundamental elements of effective interactions, enriching conversations by adding emotional and contextual depth. Recognizing the importance of non-linguistic content in communication, we present AV-EmoDialog, a dialogue system designed to exploit verbal and non-verbal information from users' audio-visual inputs to generate more responsive and empathetic interactions. AV-EmoDialog systematically exploits the emotional cues in audio-visual dialogues; extracting speech content and emotional tones from speech, analyzing fine-grained facial expressions from visuals, and integrating these cues to generate emotionally aware responses in an end-to-end manner. Through extensive experiments, we validate that the proposed AV-EmoDialog outperforms existing multimodal LLMs in generating not only emotionally appropriate but also contextually appropriate responses.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
Park, Se Jin
Kim, Yeonju
Rha, Hyeongseop
Godiva, Bella
Ro, Yong Man
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
In human communication, both verbal and non-verbal cues play a crucial role in conveying emotions, intentions, and meaning beyond words alone. These non-linguistic information, such as facial expressions, eye contact, voice tone, and pitch, are fundamental elements of effective interactions, enriching conversations by adding emotional and contextual depth. Recognizing the importance of non-linguistic content in communication, we present AV-EmoDialog, a dialogue system designed to exploit verbal and non-verbal information from users' audio-visual inputs to generate more responsive and empathetic interactions. AV-EmoDialog systematically exploits the emotional cues in audio-visual dialogues; extracting speech content and emotional tones from speech, analyzing fine-grained facial expressions from visuals, and integrating these cues to generate emotionally aware responses in an end-to-end manner. Through extensive experiments, we validate that the proposed AV-EmoDialog outperforms existing multimodal LLMs in generating not only emotionally appropriate but also contextually appropriate responses.
title AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2412.17292