FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Setyawan, Novendra, Sun, Chi-Chia, Hsu, Mao-Hsiu, Kuo, Wen-Kai, Hsieh, Jun-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917126571294720
author Setyawan, Novendra
Sun, Chi-Chia
Hsu, Mao-Hsiu
Kuo, Wen-Kai
Hsieh, Jun-Wei
author_facet Setyawan, Novendra
Sun, Chi-Chia
Hsu, Mao-Hsiu
Kuo, Wen-Kai
Hsieh, Jun-Wei
contents This paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovative and lightweight Multi-Head Linear Attention (MHLA) mechanism. By combining MHLA alongside a reparameterized token mixer, FaceLiVT effectively reduces computational complexity while preserving competitive accuracy. Extensive evaluations on challenging benchmarks; including LFW, CFP-FP, AgeDB-30, IJB-B, and IJB-C; highlight its superior performance compared to state-of-the-art lightweight models. MHLA notably improves inference speed, allowing FaceLiVT to deliver high accuracy with lower latency on mobile devices. Specifically, FaceLiVT is 8.6 faster than EdgeFace, a recent hybrid CNN-Transformer model optimized for edge devices, and 21.2 faster than a pure ViT-Based model. With its balanced design, FaceLiVT offers an efficient and practical solution for real-time face recognition on resource-constrained platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
Setyawan, Novendra
Sun, Chi-Chia
Hsu, Mao-Hsiu
Kuo, Wen-Kai
Hsieh, Jun-Wei
Computer Vision and Pattern Recognition
This paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovative and lightweight Multi-Head Linear Attention (MHLA) mechanism. By combining MHLA alongside a reparameterized token mixer, FaceLiVT effectively reduces computational complexity while preserving competitive accuracy. Extensive evaluations on challenging benchmarks; including LFW, CFP-FP, AgeDB-30, IJB-B, and IJB-C; highlight its superior performance compared to state-of-the-art lightweight models. MHLA notably improves inference speed, allowing FaceLiVT to deliver high accuracy with lower latency on mobile devices. Specifically, FaceLiVT is 8.6 faster than EdgeFace, a recent hybrid CNN-Transformer model optimized for edge devices, and 21.2 faster than a pure ViT-Based model. With its balanced design, FaceLiVT offers an efficient and practical solution for real-time face recognition on resource-constrained platforms.
title FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.10361