SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Zebang, Tu, Shuyuan, Huang, Dawei, Li, Minghan, Peng, Xiaojiang, Cheng, Zhi-Qi, Hauptmann, Alexander G.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909293336330240
author Cheng, Zebang
Tu, Shuyuan
Huang, Dawei
Li, Minghan
Peng, Xiaojiang
Cheng, Zhi-Qi
Hauptmann, Alexander G.
author_facet Cheng, Zebang
Tu, Shuyuan
Huang, Dawei
Li, Minghan
Peng, Xiaojiang
Cheng, Zhi-Qi
Hauptmann, Alexander G.
contents This paper presents our winning approach for the MER-NOISE and MER-OV tracks of the MER2024 Challenge on multimodal emotion recognition. Our system leverages the advanced emotional understanding capabilities of Emotion-LLaMA to generate high-quality annotations for unlabeled samples, addressing the challenge of limited labeled data. To enhance multimodal fusion while mitigating modality-specific noise, we introduce Conv-Attention, a lightweight and efficient hybrid framework. Extensive experimentation vali-dates the effectiveness of our approach. In the MER-NOISE track, our system achieves a state-of-the-art weighted average F-score of 85.30%, surpassing the second and third-place teams by 1.47% and 1.65%, respectively. For the MER-OV track, our utilization of Emotion-LLaMA for open-vocabulary annotation yields an 8.52% improvement in average accuracy and recall compared to GPT-4V, securing the highest score among all participating large multimodal models. The code and model for Emotion-LLaMA are available at https://github.com/ZebangCheng/Emotion-LLaMA.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10500
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
Cheng, Zebang
Tu, Shuyuan
Huang, Dawei
Li, Minghan
Peng, Xiaojiang
Cheng, Zhi-Qi
Hauptmann, Alexander G.
Multimedia
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
This paper presents our winning approach for the MER-NOISE and MER-OV tracks of the MER2024 Challenge on multimodal emotion recognition. Our system leverages the advanced emotional understanding capabilities of Emotion-LLaMA to generate high-quality annotations for unlabeled samples, addressing the challenge of limited labeled data. To enhance multimodal fusion while mitigating modality-specific noise, we introduce Conv-Attention, a lightweight and efficient hybrid framework. Extensive experimentation vali-dates the effectiveness of our approach. In the MER-NOISE track, our system achieves a state-of-the-art weighted average F-score of 85.30%, surpassing the second and third-place teams by 1.47% and 1.65%, respectively. For the MER-OV track, our utilization of Emotion-LLaMA for open-vocabulary annotation yields an 8.52% improvement in average accuracy and recall compared to GPT-4V, securing the highest score among all participating large multimodal models. The code and model for Emotion-LLaMA are available at https://github.com/ZebangCheng/Emotion-LLaMA.
title SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
topic Multimedia
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2408.10500