The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pham, Dinh Nam, Avramidis, Eleftherios
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918107309670400
author Pham, Dinh Nam
Avramidis, Eleftherios
author_facet Pham, Dinh Nam
Avramidis, Eleftherios
contents Non-manual facial features play a crucial role in sign language communication, yet their importance in automatic sign language recognition (ASLR) remains underexplored. While prior studies have shown that incorporating facial features can improve recognition, related work often relies on hand-crafted feature extraction and fails to go beyond the comparison of manual features versus the combination of manual and facial features. In this work, we systematically investigate the contribution of distinct facial regionseyes, mouth, and full faceusing two different deep learning models (a CNN-based model and a transformer-based model) trained on an SLR dataset of isolated signs with randomly selected classes. Through quantitative performance and qualitative saliency map evaluation, we reveal that the mouth is the most important non-manual facial feature, significantly improving accuracy. Our findings highlight the necessity of incorporating facial features in ASLR.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?
Pham, Dinh Nam
Avramidis, Eleftherios
Computer Vision and Pattern Recognition
Computation and Language
Image and Video Processing
Non-manual facial features play a crucial role in sign language communication, yet their importance in automatic sign language recognition (ASLR) remains underexplored. While prior studies have shown that incorporating facial features can improve recognition, related work often relies on hand-crafted feature extraction and fails to go beyond the comparison of manual features versus the combination of manual and facial features. In this work, we systematically investigate the contribution of distinct facial regionseyes, mouth, and full faceusing two different deep learning models (a CNN-based model and a transformer-based model) trained on an SLR dataset of isolated signs with randomly selected classes. Through quantitative performance and qualitative saliency map evaluation, we reveal that the mouth is the most important non-manual facial feature, significantly improving accuracy. Our findings highlight the necessity of incorporating facial features in ASLR.
title The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?
topic Computer Vision and Pattern Recognition
Computation and Language
Image and Video Processing
url https://arxiv.org/abs/2507.20884