AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Jiachen, Baevski, Alexei, Hsu, Wei-Ning, Auli, Michael
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916098952134656
author Lian, Jiachen
Baevski, Alexei
Hsu, Wei-Ning
Auli, Michael
author_facet Lian, Jiachen
Baevski, Alexei
Hsu, Wei-Ning
Auli, Michael
contents Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint representations of both modalities. In this paper, we introduce AV-data2vec which addresses these challenges and builds audio-visual representations based on predicting contextualized representations which has been successful in the uni-modal case. The model uses a shared transformer encoder for both audio and video and can combine both modalities to improve speech recognition. Results on LRS3 show that AV-data2vec consistently outperforms existing methods under all settings with the same amount of data and model size.
format Preprint
id arxiv_https___arxiv_org_abs_2302_06419
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
Lian, Jiachen
Baevski, Alexei
Hsu, Wei-Ning
Auli, Michael
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint representations of both modalities. In this paper, we introduce AV-data2vec which addresses these challenges and builds audio-visual representations based on predicting contextualized representations which has been successful in the uni-modal case. The model uses a shared transformer encoder for both audio and video and can combine both modalities to improve speech recognition. Results on LRS3 show that AV-data2vec consistently outperforms existing methods under all settings with the same amount of data and model size.
title AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2302.06419