LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Parcollet, Titouan, Nguyen, Ha, Evain, Solene, Boito, Marcely Zanon, Pupier, Adrien, Mdhaffar, Salima, Le, Hang, Alisamir, Sina, Tomashenko, Natalia, Dinarelli, Marco, Zhang, Shucong, Allauzen, Alexandre, Coavoux, Maximin, Esteve, Yannick, Rouvier, Mickael, Goulian, Jerome, Lecouteux, Benjamin, Portet, Francois, Rossato, Solange, Ringeval, Fabien, Schwab, Didier, Besacier, Laurent
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911799964598272
author Parcollet, Titouan
Nguyen, Ha
Evain, Solene
Boito, Marcely Zanon
Pupier, Adrien
Mdhaffar, Salima
Le, Hang
Alisamir, Sina
Tomashenko, Natalia
Dinarelli, Marco
Zhang, Shucong
Allauzen, Alexandre
Coavoux, Maximin
Esteve, Yannick
Rouvier, Mickael
Goulian, Jerome
Lecouteux, Benjamin
Portet, Francois
Rossato, Solange
Ringeval, Fabien
Schwab, Didier
Besacier, Laurent
author_facet Parcollet, Titouan
Nguyen, Ha
Evain, Solene
Boito, Marcely Zanon
Pupier, Adrien
Mdhaffar, Salima
Le, Hang
Alisamir, Sina
Tomashenko, Natalia
Dinarelli, Marco
Zhang, Shucong
Allauzen, Alexandre
Coavoux, Maximin
Esteve, Yannick
Rouvier, Mickael
Goulian, Jerome
Lecouteux, Benjamin
Portet, Francois
Rossato, Solange
Ringeval, Fabien
Schwab, Didier
Besacier, Laurent
contents Self-supervised learning (SSL) is at the origin of unprecedented improvements in many different domains including computer vision and natural language processing. Speech processing drastically benefitted from SSL as most of the current domain-related tasks are now being approached with pre-trained models. This work introduces LeBenchmark 2.0 an open-source framework for assessing and building SSL-equipped French speech technologies. It includes documented, large-scale and heterogeneous corpora with up to 14,000 hours of heterogeneous speech, ten pre-trained SSL wav2vec 2.0 models containing from 26 million to one billion learnable parameters shared with the community, and an evaluation protocol made of six downstream tasks to complement existing benchmarks. LeBenchmark 2.0 also presents unique perspectives on pre-trained SSL models for speech with the investigation of frozen versus fine-tuned downstream models, task-agnostic versus task-specific pre-trained models as well as a discussion on the carbon footprint of large-scale model training. Overall, the newly introduced models trained on 14,000 hours of French speech outperform multilingual and previous LeBenchmark SSL models across the benchmark but also required up to four times more energy for pre-training.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05472
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
Parcollet, Titouan
Nguyen, Ha
Evain, Solene
Boito, Marcely Zanon
Pupier, Adrien
Mdhaffar, Salima
Le, Hang
Alisamir, Sina
Tomashenko, Natalia
Dinarelli, Marco
Zhang, Shucong
Allauzen, Alexandre
Coavoux, Maximin
Esteve, Yannick
Rouvier, Mickael
Goulian, Jerome
Lecouteux, Benjamin
Portet, Francois
Rossato, Solange
Ringeval, Fabien
Schwab, Didier
Besacier, Laurent
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Self-supervised learning (SSL) is at the origin of unprecedented improvements in many different domains including computer vision and natural language processing. Speech processing drastically benefitted from SSL as most of the current domain-related tasks are now being approached with pre-trained models. This work introduces LeBenchmark 2.0 an open-source framework for assessing and building SSL-equipped French speech technologies. It includes documented, large-scale and heterogeneous corpora with up to 14,000 hours of heterogeneous speech, ten pre-trained SSL wav2vec 2.0 models containing from 26 million to one billion learnable parameters shared with the community, and an evaluation protocol made of six downstream tasks to complement existing benchmarks. LeBenchmark 2.0 also presents unique perspectives on pre-trained SSL models for speech with the investigation of frozen versus fine-tuned downstream models, task-agnostic versus task-specific pre-trained models as well as a discussion on the carbon footprint of large-scale model training. Overall, the newly introduced models trained on 14,000 hours of French speech outperform multilingual and previous LeBenchmark SSL models across the benchmark but also required up to four times more energy for pre-training.
title LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2309.05472