AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeo, Jeong Hun, Kim, Minsu, Choi, Jeongsoo, Kim, Dae Hoe, Ro, Yong Man
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916088964448256
author Yeo, Jeong Hun
Kim, Minsu
Choi, Jeongsoo
Kim, Dae Hoe
Ro, Yong Man
author_facet Yeo, Jeong Hun
Kim, Minsu
Choi, Jeongsoo
Kim, Dae Hoe
Ro, Yong Man
contents Visual Speech Recognition (VSR) is the task of predicting spoken words from silent lip movements. VSR is regarded as a challenging task because of the insufficient information on lip movements. In this paper, we propose an Audio Knowledge empowered Visual Speech Recognition framework (AKVSR) to complement the insufficient speech information of visual modality by using audio modality. Different from the previous methods, the proposed AKVSR 1) utilizes rich audio knowledge encoded by a large-scale pretrained audio model, 2) saves the linguistic information of audio knowledge in compact audio memory by discarding the non-linguistic information from the audio through quantization, and 3) includes Audio Bridging Module which can find the best-matched audio features from the compact audio memory, which makes our training possible without audio inputs, once after the compact audio memory is composed. We validate the effectiveness of the proposed method through extensive experiments, and achieve new state-of-the-art performances on the widely-used LRS3 dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2308_07593
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
Yeo, Jeong Hun
Kim, Minsu
Choi, Jeongsoo
Kim, Dae Hoe
Ro, Yong Man
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Image and Video Processing
Visual Speech Recognition (VSR) is the task of predicting spoken words from silent lip movements. VSR is regarded as a challenging task because of the insufficient information on lip movements. In this paper, we propose an Audio Knowledge empowered Visual Speech Recognition framework (AKVSR) to complement the insufficient speech information of visual modality by using audio modality. Different from the previous methods, the proposed AKVSR 1) utilizes rich audio knowledge encoded by a large-scale pretrained audio model, 2) saves the linguistic information of audio knowledge in compact audio memory by discarding the non-linguistic information from the audio through quantization, and 3) includes Audio Bridging Module which can find the best-matched audio features from the compact audio memory, which makes our training possible without audio inputs, once after the compact audio memory is composed. We validate the effectiveness of the proposed method through extensive experiments, and achieve new state-of-the-art performances on the widely-used LRS3 dataset.
title AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
topic Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2308.07593