Batch Transformer: Look for Attention in Batch

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Her, Myung Beom, Jeong, Jisu, Song, Hojoon, Han, Ji-Hyeong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908636973891584
author Her, Myung Beom
Jeong, Jisu
Song, Hojoon
Han, Ji-Hyeong
author_facet Her, Myung Beom
Jeong, Jisu
Song, Hojoon
Han, Ji-Hyeong
contents Facial expression recognition (FER) has received considerable attention in computer vision, with "in-the-wild" environments such as human-computer interaction. However, FER images contain uncertainties such as occlusion, low resolution, pose variation, illumination variation, and subjectivity, which includes some expressions that do not match the target label. Consequently, little information is obtained from a noisy single image and it is not trusted. This could significantly degrade the performance of the FER task. To address this issue, we propose a batch transformer (BT), which consists of the proposed class batch attention (CBA) module, to prevent overfitting in noisy data and extract trustworthy information by training on features reflected from several images in a batch, rather than information from a single image. We also propose multi-level attention (MLA) to prevent overfitting the specific features by capturing correlations between each level. In this paper, we present a batch transformer network (BTN) that combines the above proposals. Experimental results on various FER benchmark datasets show that the proposed BTN consistently outperforms the state-ofthe-art in FER datasets. Representative results demonstrate the promise of the proposed BTN for FER.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04218
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Batch Transformer: Look for Attention in Batch
Her, Myung Beom
Jeong, Jisu
Song, Hojoon
Han, Ji-Hyeong
Computer Vision and Pattern Recognition
Artificial Intelligence
Facial expression recognition (FER) has received considerable attention in computer vision, with "in-the-wild" environments such as human-computer interaction. However, FER images contain uncertainties such as occlusion, low resolution, pose variation, illumination variation, and subjectivity, which includes some expressions that do not match the target label. Consequently, little information is obtained from a noisy single image and it is not trusted. This could significantly degrade the performance of the FER task. To address this issue, we propose a batch transformer (BT), which consists of the proposed class batch attention (CBA) module, to prevent overfitting in noisy data and extract trustworthy information by training on features reflected from several images in a batch, rather than information from a single image. We also propose multi-level attention (MLA) to prevent overfitting the specific features by capturing correlations between each level. In this paper, we present a batch transformer network (BTN) that combines the above proposals. Experimental results on various FER benchmark datasets show that the proposed BTN consistently outperforms the state-ofthe-art in FER datasets. Representative results demonstrate the promise of the proposed BTN for FER.
title Batch Transformer: Look for Attention in Batch
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.04218