Multimodal Emotion Recognition with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hongrui, Wu, Daiqing, Li, Yangyang, Liu, Kuien, Wang, Yuhui, Zhou, Yu, Zhao, Sicheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917516667781120
author Zhang, Hongrui
Wu, Daiqing
Li, Yangyang
Liu, Kuien
Wang, Yuhui
Zhou, Yu
Zhao, Sicheng
author_facet Zhang, Hongrui
Wu, Daiqing
Li, Yangyang
Liu, Kuien
Wang, Yuhui
Zhou, Yu
Zhao, Sicheng
contents Multimodal Emotion Recognition (MER) focuses on identifying and interpreting emotions from modality-compound inputs. Closely mirroring human cognitive processes in real-world environments, MER has drawn substantial attention from both academia and industry. Recently, a paradigm shift has been unveiled in MER, from leveraging small-scale, task-specific models to Large Language Models (LLMs). We refer to the latter as the MER-with-LLMs paradigm, which offers unprecedented generality, spurring numerous empirical attempts, even alongside speculation about LLMs' potential to achieve general emotional intelligence. However, with these new opportunities come new challenges, including the scarcity of emotionally annotated data, the affective gap both within and across modalities, and the opacity of affective interpretation. To systematically review existing research and guide future exploration, this paper categorizes prior works according to their focus on addressing these challenges into three directions: Affective Data Augmentation, Multimodal Affective Representation, and Multimodal Affective Reasoning. By thoroughly tracing the development, emerging trends, and remaining issues within each direction, this paper aims to provide a clear academic map of the MER-with-LLMs paradigm and foster its structured advancement.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21239
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Emotion Recognition with Large Language Models
Zhang, Hongrui
Wu, Daiqing
Li, Yangyang
Liu, Kuien
Wang, Yuhui
Zhou, Yu
Zhao, Sicheng
Multimedia
Multimodal Emotion Recognition (MER) focuses on identifying and interpreting emotions from modality-compound inputs. Closely mirroring human cognitive processes in real-world environments, MER has drawn substantial attention from both academia and industry. Recently, a paradigm shift has been unveiled in MER, from leveraging small-scale, task-specific models to Large Language Models (LLMs). We refer to the latter as the MER-with-LLMs paradigm, which offers unprecedented generality, spurring numerous empirical attempts, even alongside speculation about LLMs' potential to achieve general emotional intelligence. However, with these new opportunities come new challenges, including the scarcity of emotionally annotated data, the affective gap both within and across modalities, and the opacity of affective interpretation. To systematically review existing research and guide future exploration, this paper categorizes prior works according to their focus on addressing these challenges into three directions: Affective Data Augmentation, Multimodal Affective Representation, and Multimodal Affective Reasoning. By thoroughly tracing the development, emerging trends, and remaining issues within each direction, this paper aims to provide a clear academic map of the MER-with-LLMs paradigm and foster its structured advancement.
title Multimodal Emotion Recognition with Large Language Models
topic Multimedia
url https://arxiv.org/abs/2605.21239