Learning Speaker-Invariant Visual Features for Lipreading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yu, Xue, Feng, Li, Shujie, Zhang, Jinrui, Yang, Shuang, Guo, Dan, Hong, Richang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913885179609088
author Li, Yu
Xue, Feng
Li, Shujie
Zhang, Jinrui
Yang, Shuang
Guo, Dan
Hong, Richang
author_facet Li, Yu
Xue, Feng
Li, Shujie
Zhang, Jinrui
Yang, Shuang
Guo, Dan
Hong, Richang
contents Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape, color, texture), which introduce spurious correlations between vision and text. These correlations lead to suboptimal lipreading accuracy and restrict model generalization. To address this challenge, we introduce SIFLip, a speaker-invariant visual feature learning framework that disentangles speaker-specific attributes using two complementary disentanglement modules (Implicit Disentanglement and Explicit Disentanglement) to improve generalization. Specifically, since different speakers exhibit semantic consistency between lip movements and phonetic text when pronouncing the same words, our implicit disentanglement module leverages stable text embeddings as supervisory signals to learn common visual representations across speakers, implicitly decoupling speaker-specific features. Additionally, we design a speaker recognition sub-task within the main lipreading pipeline to filter speaker-specific features, then further explicitly disentangle these personalized visual features from the backbone network via gradient reversal. Experimental results demonstrate that SIFLip significantly enhances generalization performance across multiple public datasets. Experimental results demonstrate that SIFLip significantly improves generalization performance across multiple public datasets, outperforming state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Speaker-Invariant Visual Features for Lipreading
Li, Yu
Xue, Feng
Li, Shujie
Zhang, Jinrui
Yang, Shuang
Guo, Dan
Hong, Richang
Computer Vision and Pattern Recognition
Computation and Language
Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape, color, texture), which introduce spurious correlations between vision and text. These correlations lead to suboptimal lipreading accuracy and restrict model generalization. To address this challenge, we introduce SIFLip, a speaker-invariant visual feature learning framework that disentangles speaker-specific attributes using two complementary disentanglement modules (Implicit Disentanglement and Explicit Disentanglement) to improve generalization. Specifically, since different speakers exhibit semantic consistency between lip movements and phonetic text when pronouncing the same words, our implicit disentanglement module leverages stable text embeddings as supervisory signals to learn common visual representations across speakers, implicitly decoupling speaker-specific features. Additionally, we design a speaker recognition sub-task within the main lipreading pipeline to filter speaker-specific features, then further explicitly disentangle these personalized visual features from the backbone network via gradient reversal. Experimental results demonstrate that SIFLip significantly enhances generalization performance across multiple public datasets. Experimental results demonstrate that SIFLip significantly improves generalization performance across multiple public datasets, outperforming state-of-the-art methods.
title Learning Speaker-Invariant Visual Features for Lipreading
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.07572