ML-SUPERB: Multilingual Speech Universal PERformance Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Jiatong, Berrebbi, Dan, Chen, William, Chung, Ho-Lam, Hu, En-Pei, Huang, Wei Ping, Chang, Xuankai, Li, Shang-Wen, Mohamed, Abdelrahman, Lee, Hung-yi, Watanabe, Shinji
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917934678409216
author Shi, Jiatong
Berrebbi, Dan
Chen, William
Chung, Ho-Lam
Hu, En-Pei
Huang, Wei Ping
Chang, Xuankai
Li, Shang-Wen
Mohamed, Abdelrahman
Lee, Hung-yi
Watanabe, Shinji
author_facet Shi, Jiatong
Berrebbi, Dan
Chen, William
Chung, Ho-Lam
Hu, En-Pei
Huang, Wei Ping
Chang, Xuankai
Li, Shang-Wen
Mohamed, Abdelrahman
Lee, Hung-yi
Watanabe, Shinji
contents Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers English speech in its evaluation. This paper presents multilingual SUPERB (ML-SUPERB), covering 143 languages (ranging from high-resource to endangered), and considering both automatic speech recognition and language identification. Following the concept of SUPERB, ML-SUPERB utilizes frozen SSL features and employs a simple framework for multilingual tasks by learning a shallow downstream model. Similar to the SUPERB benchmark, we find speech SSL models can significantly improve performance compared to FBANK features. Furthermore, we find that multilingual models do not always perform better than their monolingual counterparts. We will release ML-SUPERB as a challenge with organized datasets and reproducible training scripts for future multilingual representation research.
format Preprint
id arxiv_https___arxiv_org_abs_2305_10615
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
Shi, Jiatong
Berrebbi, Dan
Chen, William
Chung, Ho-Lam
Hu, En-Pei
Huang, Wei Ping
Chang, Xuankai
Li, Shang-Wen
Mohamed, Abdelrahman
Lee, Hung-yi
Watanabe, Shinji
Sound
Computation and Language
Audio and Speech Processing
Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers English speech in its evaluation. This paper presents multilingual SUPERB (ML-SUPERB), covering 143 languages (ranging from high-resource to endangered), and considering both automatic speech recognition and language identification. Following the concept of SUPERB, ML-SUPERB utilizes frozen SSL features and employs a simple framework for multilingual tasks by learning a shallow downstream model. Similar to the SUPERB benchmark, we find speech SSL models can significantly improve performance compared to FBANK features. Furthermore, we find that multilingual models do not always perform better than their monolingual counterparts. We will release ML-SUPERB as a challenge with organized datasets and reproducible training scripts for future multilingual representation research.
title ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2305.10615