Lightweight Operations for Visual Speech Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Panagos, Iason Ioannis, Sfikas, Giorgos, Nikou, Christophoros
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929702286917632
author Panagos, Iason Ioannis
Sfikas, Giorgos
Nikou, Christophoros
author_facet Panagos, Iason Ioannis
Sfikas, Giorgos
Nikou, Christophoros
contents Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that demand powerful hardware, limiting VSR deployment on resource-constrained devices. This work addresses this limitation by developing lightweight VSR architectures. Leveraging efficient operation design paradigms, we create compact yet powerful models with reduced resource requirements and minimal accuracy loss. We train and evaluate our models on a large-scale public dataset for recognition of words from video sequences, demonstrating their effectiveness for practical applications. We also conduct an extensive array of ablative experiments to thoroughly analyze the size and complexity of each model. Code and trained models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lightweight Operations for Visual Speech Recognition
Panagos, Iason Ioannis
Sfikas, Giorgos
Nikou, Christophoros
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that demand powerful hardware, limiting VSR deployment on resource-constrained devices. This work addresses this limitation by developing lightweight VSR architectures. Leveraging efficient operation design paradigms, we create compact yet powerful models with reduced resource requirements and minimal accuracy loss. We train and evaluate our models on a large-scale public dataset for recognition of words from video sequences, demonstrating their effectiveness for practical applications. We also conduct an extensive array of ablative experiments to thoroughly analyze the size and complexity of each model. Code and trained models will be made publicly available.
title Lightweight Operations for Visual Speech Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.04834