TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yun, Taeyang, Lim, Hyunkuk, Lee, Jeonghwan, Song, Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911819925291008
author Yun, Taeyang
Lim, Hyunkuk
Lee, Jeonghwan
Song, Min
author_facet Yun, Taeyang
Lim, Hyunkuk
Lee, Jeonghwan
Song, Min
contents Emotion Recognition in Conversation (ERC) plays a crucial role in enabling dialogue systems to effectively respond to user requests. The emotions in a conversation can be identified by the representations from various modalities, such as audio, visual, and text. However, due to the weak contribution of non-verbal modalities to recognize emotions, multimodal ERC has always been considered a challenging task. In this paper, we propose Teacher-leading Multimodal fusion network for ERC (TelME). TelME incorporates cross-modal knowledge distillation to transfer information from a language model acting as the teacher to the non-verbal students, thereby optimizing the efficacy of the weak modalities. We then combine multimodal features using a shifting fusion approach in which student networks support the teacher. TelME achieves state-of-the-art performance in MELD, a multi-speaker conversation dataset for ERC. Finally, we demonstrate the effectiveness of our components through additional experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12987
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
Yun, Taeyang
Lim, Hyunkuk
Lee, Jeonghwan
Song, Min
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
Emotion Recognition in Conversation (ERC) plays a crucial role in enabling dialogue systems to effectively respond to user requests. The emotions in a conversation can be identified by the representations from various modalities, such as audio, visual, and text. However, due to the weak contribution of non-verbal modalities to recognize emotions, multimodal ERC has always been considered a challenging task. In this paper, we propose Teacher-leading Multimodal fusion network for ERC (TelME). TelME incorporates cross-modal knowledge distillation to transfer information from a language model acting as the teacher to the non-verbal students, thereby optimizing the efficacy of the weak modalities. We then combine multimodal features using a shifting fusion approach in which student networks support the teacher. TelME achieves state-of-the-art performance in MELD, a multi-speaker conversation dataset for ERC. Finally, we demonstrate the effectiveness of our components through additional experiments.
title TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.12987