Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Armando, Li, Keqin, Wu, Tong, Zhao, Peng, Hong, Bo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909184846462976
author Zhu, Armando
Li, Keqin
Wu, Tong
Zhao, Peng
Hong, Bo
author_facet Zhu, Armando
Li, Keqin
Wu, Tong
Zhao, Peng
Hong, Bo
contents With wearing masks becoming a new cultural norm, facial expression recognition (FER) while taking masks into account has become a significant challenge. In this paper, we propose a unified multi-branch vision transformer for facial expression recognition and mask wearing classification tasks. Our approach extracts shared features for both tasks using a dual-branch architecture that obtains multi-scale feature representations. Furthermore, we propose a cross-task fusion phase that processes tokens for each task with separate branches, while exchanging information using a cross attention module. Our proposed framework reduces the overall complexity compared with using separate networks for both tasks by the simple yet effective cross-task fusion phase. Extensive experiments demonstrate that our proposed model performs better than or on par with different state-of-the-art methods on both facial expression recognition and facial mask wearing classification task.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14606
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
Zhu, Armando
Li, Keqin
Wu, Tong
Zhao, Peng
Hong, Bo
Computer Vision and Pattern Recognition
Artificial Intelligence
With wearing masks becoming a new cultural norm, facial expression recognition (FER) while taking masks into account has become a significant challenge. In this paper, we propose a unified multi-branch vision transformer for facial expression recognition and mask wearing classification tasks. Our approach extracts shared features for both tasks using a dual-branch architecture that obtains multi-scale feature representations. Furthermore, we propose a cross-task fusion phase that processes tokens for each task with separate branches, while exchanging information using a cross attention module. Our proposed framework reduces the overall complexity compared with using separate networks for both tasks by the simple yet effective cross-task fusion phase. Extensive experiments demonstrate that our proposed model performs better than or on par with different state-of-the-art methods on both facial expression recognition and facial mask wearing classification task.
title Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2404.14606