EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xing, Bohao, Yu, Zitong, Liu, Xin, Yuan, Kaishen, Ye, Qilang, Xie, Weicheng, Yue, Huanjing, Yang, Jingyu, Kälviäinen, Heikki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911997658923008
author Xing, Bohao
Yu, Zitong
Liu, Xin
Yuan, Kaishen
Ye, Qilang
Xie, Weicheng
Yue, Huanjing
Yang, Jingyu
Kälviäinen, Heikki
author_facet Xing, Bohao
Yu, Zitong
Liu, Xin
Yuan, Kaishen
Ye, Qilang
Xie, Weicheng
Yue, Huanjing
Yang, Jingyu
Kälviäinen, Heikki
contents Facial expression recognition (FER) is an important research topic in emotional artificial intelligence. In recent decades, researchers have made remarkable progress. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework, making their application in multimodal emotion understanding and human-computer interaction difficult. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. However, directly applying pre-trained MLLMs to FER still faces several challenges. Our zero-shot evaluations of existing open-source MLLMs on FER indicate a significant performance gap compared to GPT-4V and current supervised state-of-the-art (SOTA) methods. In this paper, we aim to enhance MLLMs' capabilities in understanding facial expressions. We first generate instruction data for five FER datasets with Gemini. We then propose a novel MLLM, named EMO-LLaMA, which incorporates facial priors from a pretrained facial analysis network to enhance human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Additionally, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across different human groups. Extensive experiments show that EMO-LLaMA achieves SOTA-comparable or competitive results across both static and dynamic FER datasets. The instruction dataset and code are available at https://github.com/xxtars/EMO-LLaMA.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
Xing, Bohao
Yu, Zitong
Liu, Xin
Yuan, Kaishen
Ye, Qilang
Xie, Weicheng
Yue, Huanjing
Yang, Jingyu
Kälviäinen, Heikki
Computer Vision and Pattern Recognition
Facial expression recognition (FER) is an important research topic in emotional artificial intelligence. In recent decades, researchers have made remarkable progress. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework, making their application in multimodal emotion understanding and human-computer interaction difficult. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. However, directly applying pre-trained MLLMs to FER still faces several challenges. Our zero-shot evaluations of existing open-source MLLMs on FER indicate a significant performance gap compared to GPT-4V and current supervised state-of-the-art (SOTA) methods. In this paper, we aim to enhance MLLMs' capabilities in understanding facial expressions. We first generate instruction data for five FER datasets with Gemini. We then propose a novel MLLM, named EMO-LLaMA, which incorporates facial priors from a pretrained facial analysis network to enhance human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Additionally, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across different human groups. Extensive experiments show that EMO-LLaMA achieves SOTA-comparable or competitive results across both static and dynamic FER datasets. The instruction dataset and code are available at https://github.com/xxtars/EMO-LLaMA.
title EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.11424