BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haliassos, Alexandros, Zinonos, Andreas, Mira, Rodrigo, Petridis, Stavros, Pantic, Maja
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916189638230016
author Haliassos, Alexandros
Zinonos, Andreas
Mira, Rodrigo
Petridis, Stavros
Pantic, Maja
author_facet Haliassos, Alexandros
Zinonos, Andreas
Mira, Rodrigo
Petridis, Stavros
Pantic, Maja
contents Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations entirely from raw audio-visual data. Our modifications to RAVEn enable BRAVEn to achieve state-of-the-art results among self-supervised methods in various settings. Moreover, we observe favourable scaling behaviour by increasing the amount of unlabelled data well beyond other self-supervised works. In particular, we achieve 20.0% / 1.7% word error rate for VSR / ASR on the LRS3 test set, with only 30 hours of labelled data and no external ASR models. Our results suggest that readily available unlabelled audio-visual data can largely replace costly transcribed data.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02098
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
Haliassos, Alexandros
Zinonos, Andreas
Mira, Rodrigo
Petridis, Stavros
Pantic, Maja
Computer Vision and Pattern Recognition
Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations entirely from raw audio-visual data. Our modifications to RAVEn enable BRAVEn to achieve state-of-the-art results among self-supervised methods in various settings. Moreover, we observe favourable scaling behaviour by increasing the amount of unlabelled data well beyond other self-supervised works. In particular, we achieve 20.0% / 1.7% word error rate for VSR / ASR on the LRS3 test set, with only 30 hours of labelled data and no external ASR models. Our results suggest that readily available unlabelled audio-visual data can largely replace costly transcribed data.
title BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.02098