Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Linzhi, Zhang, Xingyu, Zhang, Yakun, Zheng, Changyan, Liu, Tiejun, Xie, Liang, Yan, Ye, Yin, Erwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909187355705344
author Wu, Linzhi
Zhang, Xingyu
Zhang, Yakun
Zheng, Changyan
Liu, Tiejun
Xie, Liang
Yan, Ye
Yin, Erwei
author_facet Wu, Linzhi
Zhang, Xingyu
Zhang, Yakun
Zheng, Changyan
Liu, Tiejun
Xie, Liang
Yan, Ye
Yin, Erwei
contents Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity changes, poses a challenging problem due to inter-speaker variability. A well-trained lip reading system may perform poorly when handling a brand new speaker. To learn a speaker-robust lip reading model, a key insight is to reduce visual variations across speakers, avoiding the model overfitting to specific speakers. In this work, in view of both input visual clues and latent representations based on a hybrid CTC/attention architecture, we propose to exploit the lip landmark-guided fine-grained visual clues instead of frequently-used mouth-cropped images as input features, diminishing speaker-specific appearance characteristics. Furthermore, a max-min mutual information regularization approach is proposed to capture speaker-insensitive latent representations. Experimental evaluations on public lip reading datasets demonstrate the effectiveness of the proposed approach under the intra-speaker and inter-speaker conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
Wu, Linzhi
Zhang, Xingyu
Zhang, Yakun
Zheng, Changyan
Liu, Tiejun
Xie, Liang
Yan, Ye
Yin, Erwei
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity changes, poses a challenging problem due to inter-speaker variability. A well-trained lip reading system may perform poorly when handling a brand new speaker. To learn a speaker-robust lip reading model, a key insight is to reduce visual variations across speakers, avoiding the model overfitting to specific speakers. In this work, in view of both input visual clues and latent representations based on a hybrid CTC/attention architecture, we propose to exploit the lip landmark-guided fine-grained visual clues instead of frequently-used mouth-cropped images as input features, diminishing speaker-specific appearance characteristics. Furthermore, a max-min mutual information regularization approach is proposed to capture speaker-insensitive latent representations. Experimental evaluations on public lip reading datasets demonstrate the effectiveness of the proposed approach under the intra-speaker and inter-speaker conditions.
title Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2403.16071