AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916098952134656 |
|---|---|
| author | Lian, Jiachen Baevski, Alexei Hsu, Wei-Ning Auli, Michael |
| author_facet | Lian, Jiachen Baevski, Alexei Hsu, Wei-Ning Auli, Michael |
| contents | Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint representations of both modalities. In this paper, we introduce AV-data2vec which addresses these challenges and builds audio-visual representations based on predicting contextualized representations which has been successful in the uni-modal case. The model uses a shared transformer encoder for both audio and video and can combine both modalities to improve speech recognition. Results on LRS3 show that AV-data2vec consistently outperforms existing methods under all settings with the same amount of data and model size. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2302_06419 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations Lian, Jiachen Baevski, Alexei Hsu, Wei-Ning Auli, Michael Audio and Speech Processing Artificial Intelligence Computation and Language Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint representations of both modalities. In this paper, we introduce AV-data2vec which addresses these challenges and builds audio-visual representations based on predicting contextualized representations which has been successful in the uni-modal case. The model uses a shared transformer encoder for both audio and video and can combine both modalities to improve speech recognition. Results on LRS3 show that AV-data2vec consistently outperforms existing methods under all settings with the same amount of data and model size. |
| title | AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2302.06419 |