Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, He, Guo, Pengcheng, Wan, Xucheng, Zhou, Huan, Xie, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917653728198656
author Wang, He
Guo, Pengcheng
Wan, Xucheng
Zhou, Huan
Xie, Lei
author_facet Wang, He
Guo, Pengcheng
Wan, Xucheng
Zhou, Huan
Xie, Lei
contents Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single scale. In this paper, we propose to enhance lip-reading by incorporating multi-scale video data and multi-encoder. Specifically, we first propose a novel multi-scale lip motion extraction algorithm based on the size of the speaker's face and an Enhanced ResNet3D visual front-end (VFE) to extract lip features at different scales. For the multi-encoder, in addition to the mainstream Transformer and Conformer, we also incorporate the recently proposed Branchformer and E-Branchformer as visual encoders. In the experiments, we explore the influence of different video data scales and encoders on ALR system performance and fuse the texts transcribed by all ALR systems using recognizer output voting error reduction (ROVER). Finally, our proposed approach placed second in the ICME 2024 ChatCLR Challenge Task 2, with a 21.52% reduction in character error rate (CER) compared to the official baseline on the evaluation set.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05466
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder
Wang, He
Guo, Pengcheng
Wan, Xucheng
Zhou, Huan
Xie, Lei
Computer Vision and Pattern Recognition
Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single scale. In this paper, we propose to enhance lip-reading by incorporating multi-scale video data and multi-encoder. Specifically, we first propose a novel multi-scale lip motion extraction algorithm based on the size of the speaker's face and an Enhanced ResNet3D visual front-end (VFE) to extract lip features at different scales. For the multi-encoder, in addition to the mainstream Transformer and Conformer, we also incorporate the recently proposed Branchformer and E-Branchformer as visual encoders. In the experiments, we explore the influence of different video data scales and encoders on ALR system performance and fuse the texts transcribed by all ALR systems using recognizer output voting error reduction (ROVER). Finally, our proposed approach placed second in the ICME 2024 ChatCLR Challenge Task 2, with a 21.52% reduction in character error rate (CER) compared to the official baseline on the evaluation set.
title Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.05466