Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913492180664320 |
|---|---|
| author | Saadi, Ibtissam Cunningham, Douglas W. Abdelmalik, Taleb-ahmed Hadid, Abdenour Hillali, Yassin El |
| author_facet | Saadi, Ibtissam Cunningham, Douglas W. Abdelmalik, Taleb-ahmed Hadid, Abdenour Hillali, Yassin El |
| contents | Existing methods for driver facial expression recognition (DFER) are often computationally intensive, rendering them unsuitable for real-time applications. In this work, we introduce a novel transfer learning-based dual architecture, named ShuffViT-DFER, which elegantly combines computational efficiency and accuracy. This is achieved by harnessing the strengths of two lightweight and efficient models using convolutional neural network (CNN) and vision transformers (ViT). We efficiently fuse the extracted features to enhance the performance of the model in accurately recognizing the facial expressions of the driver. Our experimental results on two benchmarking and public datasets, KMU-FED and KDEF, highlight the validity of our proposed method for real-time application with superior performance when compared to state-of-the-art methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_03438 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression Saadi, Ibtissam Cunningham, Douglas W. Abdelmalik, Taleb-ahmed Hadid, Abdenour Hillali, Yassin El Computer Vision and Pattern Recognition Existing methods for driver facial expression recognition (DFER) are often computationally intensive, rendering them unsuitable for real-time applications. In this work, we introduce a novel transfer learning-based dual architecture, named ShuffViT-DFER, which elegantly combines computational efficiency and accuracy. This is achieved by harnessing the strengths of two lightweight and efficient models using convolutional neural network (CNN) and vision transformers (ViT). We efficiently fuse the extracted features to enhance the performance of the model in accurately recognizing the facial expressions of the driver. Our experimental results on two benchmarking and public datasets, KMU-FED and KDEF, highlight the validity of our proposed method for real-time application with superior performance when compared to state-of-the-art methods. |
| title | Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.03438 |