EmMixformer: Mix transformer for eye movement recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Huafeng, Zhu, Hongyu, Jin, Xin, Song, Qun, El-Yacoubi, Mounim A., Gao, Xinbo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929338052509696
author Qin, Huafeng
Zhu, Hongyu
Jin, Xin
Song, Qun
El-Yacoubi, Mounim A.
Gao, Xinbo
author_facet Qin, Huafeng
Zhu, Hongyu
Jin, Xin
Song, Qun
El-Yacoubi, Mounim A.
Gao, Xinbo
contents Eye movement (EM) is a new highly secure biometric behavioral modality that has received increasing attention in recent years. Although deep neural networks, such as convolutional neural network (CNN), have recently achieved promising performance, current solutions fail to capture local and global temporal dependencies within eye movement data. To overcome this problem, we propose in this paper a mixed transformer termed EmMixformer to extract time and frequency domain information for eye movement recognition. To this end, we propose a mixed block consisting of three modules, transformer, attention Long short-term memory (attention LSTM), and Fourier transformer. We are the first to attempt leveraging transformer to learn long temporal dependencies within eye movement. Second, we incorporate the attention mechanism into LSTM to propose attention LSTM with the aim to learn short temporal dependencies. Third, we perform self attention in the frequency domain to learn global features. As the three modules provide complementary feature representations in terms of local and global dependencies, the proposed EmMixformer is capable of improving recognition accuracy. The experimental results on our eye movement dataset and two public eye movement datasets show that the proposed EmMixformer outperforms the state of the art by achieving the lowest verification error.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmMixformer: Mix transformer for eye movement recognition
Qin, Huafeng
Zhu, Hongyu
Jin, Xin
Song, Qun
El-Yacoubi, Mounim A.
Gao, Xinbo
Computer Vision and Pattern Recognition
Cryptography and Security
Eye movement (EM) is a new highly secure biometric behavioral modality that has received increasing attention in recent years. Although deep neural networks, such as convolutional neural network (CNN), have recently achieved promising performance, current solutions fail to capture local and global temporal dependencies within eye movement data. To overcome this problem, we propose in this paper a mixed transformer termed EmMixformer to extract time and frequency domain information for eye movement recognition. To this end, we propose a mixed block consisting of three modules, transformer, attention Long short-term memory (attention LSTM), and Fourier transformer. We are the first to attempt leveraging transformer to learn long temporal dependencies within eye movement. Second, we incorporate the attention mechanism into LSTM to propose attention LSTM with the aim to learn short temporal dependencies. Third, we perform self attention in the frequency domain to learn global features. As the three modules provide complementary feature representations in terms of local and global dependencies, the proposed EmMixformer is capable of improving recognition accuracy. The experimental results on our eye movement dataset and two public eye movement datasets show that the proposed EmMixformer outperforms the state of the art by achieving the lowest verification error.
title EmMixformer: Mix transformer for eye movement recognition
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2401.04956