InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Junxiao, Liu, Xiaozhen, Wu, Xuecheng, Yu, Fei, Wang, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913974688153600
author Xue, Junxiao
Liu, Xiaozhen
Wu, Xuecheng
Yu, Fei
Wang, Jun
author_facet Xue, Junxiao
Liu, Xiaozhen
Wu, Xuecheng
Yu, Fei
Wang, Jun
contents Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to variability across sequences and the uneven distribution of information within each sequence. To tackle this, we introduce InfoSyncNet, a non-uniform sequence modeling network enhanced by tailored data augmentation techniques. Central to InfoSyncNet is a non-uniform quantization module positioned between the encoder and decoder, enabling dynamic adjustment to the network's focus and effectively handling the natural inconsistencies in visual speech data. Additionally, multiple training strategies are incorporated to enhance the model's capability to handle variations in lighting and the speaker's orientation. Comprehensive experiments on the LRW and LRW1000 datasets confirm the superiority of InfoSyncNet, achieving new state-of-the-art accuracies of 92.0% and 60.7% Top-1 ACC. The code is available for download (see comments).
format Preprint
id arxiv_https___arxiv_org_abs_2508_02460
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
Xue, Junxiao
Liu, Xiaozhen
Wu, Xuecheng
Yu, Fei
Wang, Jun
Computer Vision and Pattern Recognition
Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to variability across sequences and the uneven distribution of information within each sequence. To tackle this, we introduce InfoSyncNet, a non-uniform sequence modeling network enhanced by tailored data augmentation techniques. Central to InfoSyncNet is a non-uniform quantization module positioned between the encoder and decoder, enabling dynamic adjustment to the network's focus and effectively handling the natural inconsistencies in visual speech data. Additionally, multiple training strategies are incorporated to enhance the model's capability to handle variations in lighting and the speaker's orientation. Comprehensive experiments on the LRW and LRW1000 datasets confirm the superiority of InfoSyncNet, achieving new state-of-the-art accuracies of 92.0% and 60.7% Top-1 ACC. The code is available for download (see comments).
title InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.02460