FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917126571294720 |
|---|---|
| author | Setyawan, Novendra Sun, Chi-Chia Hsu, Mao-Hsiu Kuo, Wen-Kai Hsieh, Jun-Wei |
| author_facet | Setyawan, Novendra Sun, Chi-Chia Hsu, Mao-Hsiu Kuo, Wen-Kai Hsieh, Jun-Wei |
| contents | This paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovative and lightweight Multi-Head Linear Attention (MHLA) mechanism. By combining MHLA alongside a reparameterized token mixer, FaceLiVT effectively reduces computational complexity while preserving competitive accuracy. Extensive evaluations on challenging benchmarks; including LFW, CFP-FP, AgeDB-30, IJB-B, and IJB-C; highlight its superior performance compared to state-of-the-art lightweight models. MHLA notably improves inference speed, allowing FaceLiVT to deliver high accuracy with lower latency on mobile devices. Specifically, FaceLiVT is 8.6 faster than EdgeFace, a recent hybrid CNN-Transformer model optimized for edge devices, and 21.2 faster than a pure ViT-Based model. With its balanced design, FaceLiVT offers an efficient and practical solution for real-time face recognition on resource-constrained platforms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_10361 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device Setyawan, Novendra Sun, Chi-Chia Hsu, Mao-Hsiu Kuo, Wen-Kai Hsieh, Jun-Wei Computer Vision and Pattern Recognition This paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovative and lightweight Multi-Head Linear Attention (MHLA) mechanism. By combining MHLA alongside a reparameterized token mixer, FaceLiVT effectively reduces computational complexity while preserving competitive accuracy. Extensive evaluations on challenging benchmarks; including LFW, CFP-FP, AgeDB-30, IJB-B, and IJB-C; highlight its superior performance compared to state-of-the-art lightweight models. MHLA notably improves inference speed, allowing FaceLiVT to deliver high accuracy with lower latency on mobile devices. Specifically, FaceLiVT is 8.6 faster than EdgeFace, a recent hybrid CNN-Transformer model optimized for edge devices, and 21.2 faster than a pure ViT-Based model. With its balanced design, FaceLiVT offers an efficient and practical solution for real-time face recognition on resource-constrained platforms. |
| title | FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.10361 |