AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Zheng, Chen, Haoyu, Chen, Lan, Sun, Haiyang, Sun, Licai, Ren, Yong, Cheng, Zebang, Liu, Bin, Liu, Rui, Peng, Xiaojiang, Yi, Jiangyan, Tao, Jianhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913824961986560
author Lian, Zheng
Chen, Haoyu
Chen, Lan
Sun, Haiyang
Sun, Licai
Ren, Yong
Cheng, Zebang
Liu, Bin
Liu, Rui
Peng, Xiaojiang
Yi, Jiangyan
Tao, Jianhua
author_facet Lian, Zheng
Chen, Haoyu
Chen, Lan
Sun, Haiyang
Sun, Licai
Ren, Yong
Cheng, Zebang
Liu, Bin
Liu, Rui
Peng, Xiaojiang
Yi, Jiangyan
Tao, Jianhua
contents The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level, from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16566
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
Lian, Zheng
Chen, Haoyu
Chen, Lan
Sun, Haiyang
Sun, Licai
Ren, Yong
Cheng, Zebang
Liu, Bin
Liu, Rui
Peng, Xiaojiang
Yi, Jiangyan
Tao, Jianhua
Human-Computer Interaction
The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level, from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
title AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2501.16566