Adaptive-avg-pooling based Attention Vision Transformer for Face Anti-spoofing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Jichen, Chen, Fangfan, Das, Rohan Kumar, Zhu, Zhengyu, Zhang, Shunsi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909067560091648
author Yang, Jichen
Chen, Fangfan
Das, Rohan Kumar
Zhu, Zhengyu
Zhang, Shunsi
author_facet Yang, Jichen
Chen, Fangfan
Das, Rohan Kumar
Zhu, Zhengyu
Zhang, Shunsi
contents Traditional vision transformer consists of two parts: transformer encoder and multi-layer perception (MLP). The former plays the role of feature learning to obtain better representation, while the latter plays the role of classification. Here, the MLP is constituted of two fully connected (FC) layers, average value computing, FC layer and softmax layer. However, due to the use of average value computing module, some useful information may get lost, which we plan to preserve by the use of alternative framework. In this work, we propose a novel vision transformer referred to as adaptive-avg-pooling based attention vision transformer (AAViT) that uses modules of adaptive average pooling and attention to replace the module of average value computing. We explore the proposed AAViT for the studies on face anti-spoofing using Replay-Attack database. The experiments show that the AAViT outperforms vision transformer in face anti-spoofing by producing a reduced equal error rate. In addition, we found that the proposed AAViT can perform much better than some commonly used neural networks such as ResNet and some other known systems on the Replay-Attack corpus.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04953
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive-avg-pooling based Attention Vision Transformer for Face Anti-spoofing
Yang, Jichen
Chen, Fangfan
Das, Rohan Kumar
Zhu, Zhengyu
Zhang, Shunsi
Image and Video Processing
Signal Processing
Traditional vision transformer consists of two parts: transformer encoder and multi-layer perception (MLP). The former plays the role of feature learning to obtain better representation, while the latter plays the role of classification. Here, the MLP is constituted of two fully connected (FC) layers, average value computing, FC layer and softmax layer. However, due to the use of average value computing module, some useful information may get lost, which we plan to preserve by the use of alternative framework. In this work, we propose a novel vision transformer referred to as adaptive-avg-pooling based attention vision transformer (AAViT) that uses modules of adaptive average pooling and attention to replace the module of average value computing. We explore the proposed AAViT for the studies on face anti-spoofing using Replay-Attack database. The experiments show that the AAViT outperforms vision transformer in face anti-spoofing by producing a reduced equal error rate. In addition, we found that the proposed AAViT can perform much better than some commonly used neural networks such as ResNet and some other known systems on the Replay-Attack corpus.
title Adaptive-avg-pooling based Attention Vision Transformer for Face Anti-spoofing
topic Image and Video Processing
Signal Processing
url https://arxiv.org/abs/2401.04953