Emotic Masked Autoencoder with Attention Fusion for Facial Expression Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen-Xuan, Bach, Nguyen-Hoang, Thien, Nguyen, Thanh-Huy, Tai-Do, Nhu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911874033909760
author Nguyen-Xuan, Bach
Nguyen-Hoang, Thien
Nguyen, Thanh-Huy
Tai-Do, Nhu
author_facet Nguyen-Xuan, Bach
Nguyen-Hoang, Thien
Nguyen, Thanh-Huy
Tai-Do, Nhu
contents Facial Expression Recognition (FER) is a critical task within computer vision with diverse applications across various domains. Addressing the challenge of limited FER datasets, which hampers the generalization capability of expression recognition models, is imperative for enhancing performance. Our paper presents an innovative approach integrating the MAE-Face self-supervised learning (SSL) method and multi-view Fusion Attention mechanism for expression classification, particularly showcased in the 6th Affective Behavior Analysis in-the-wild (ABAW) competition. By utilizing low-level feature information from the ipsilateral view (auxiliary view) before learning the high-level feature that emphasizes the shift in the human facial expression, our work seeks to provide a straightforward yet innovative way to improve the examined view (main view). We also suggest easy-to-implement and no-training frameworks aimed at highlighting key facial features to determine if such features can serve as guides for the model, focusing on pivotal local elements. The efficacy of this method is validated by improvements in model performance on the Aff-wild2 dataset, as observed in both training and validation contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Emotic Masked Autoencoder with Attention Fusion for Facial Expression Recognition
Nguyen-Xuan, Bach
Nguyen-Hoang, Thien
Nguyen, Thanh-Huy
Tai-Do, Nhu
Computer Vision and Pattern Recognition
Facial Expression Recognition (FER) is a critical task within computer vision with diverse applications across various domains. Addressing the challenge of limited FER datasets, which hampers the generalization capability of expression recognition models, is imperative for enhancing performance. Our paper presents an innovative approach integrating the MAE-Face self-supervised learning (SSL) method and multi-view Fusion Attention mechanism for expression classification, particularly showcased in the 6th Affective Behavior Analysis in-the-wild (ABAW) competition. By utilizing low-level feature information from the ipsilateral view (auxiliary view) before learning the high-level feature that emphasizes the shift in the human facial expression, our work seeks to provide a straightforward yet innovative way to improve the examined view (main view). We also suggest easy-to-implement and no-training frameworks aimed at highlighting key facial features to determine if such features can serve as guides for the model, focusing on pivotal local elements. The efficacy of this method is validated by improvements in model performance on the Aff-wild2 dataset, as observed in both training and validation contexts.
title Emotic Masked Autoencoder with Attention Fusion for Facial Expression Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.13039