ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917934678409216 |
|---|---|
| author | Shi, Jiatong Berrebbi, Dan Chen, William Chung, Ho-Lam Hu, En-Pei Huang, Wei Ping Chang, Xuankai Li, Shang-Wen Mohamed, Abdelrahman Lee, Hung-yi Watanabe, Shinji |
| author_facet | Shi, Jiatong Berrebbi, Dan Chen, William Chung, Ho-Lam Hu, En-Pei Huang, Wei Ping Chang, Xuankai Li, Shang-Wen Mohamed, Abdelrahman Lee, Hung-yi Watanabe, Shinji |
| contents | Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers English speech in its evaluation. This paper presents multilingual SUPERB (ML-SUPERB), covering 143 languages (ranging from high-resource to endangered), and considering both automatic speech recognition and language identification. Following the concept of SUPERB, ML-SUPERB utilizes frozen SSL features and employs a simple framework for multilingual tasks by learning a shallow downstream model. Similar to the SUPERB benchmark, we find speech SSL models can significantly improve performance compared to FBANK features. Furthermore, we find that multilingual models do not always perform better than their monolingual counterparts. We will release ML-SUPERB as a challenge with organized datasets and reproducible training scripts for future multilingual representation research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_10615 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ML-SUPERB: Multilingual Speech Universal PERformance Benchmark Shi, Jiatong Berrebbi, Dan Chen, William Chung, Ho-Lam Hu, En-Pei Huang, Wei Ping Chang, Xuankai Li, Shang-Wen Mohamed, Abdelrahman Lee, Hung-yi Watanabe, Shinji Sound Computation and Language Audio and Speech Processing Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers English speech in its evaluation. This paper presents multilingual SUPERB (ML-SUPERB), covering 143 languages (ranging from high-resource to endangered), and considering both automatic speech recognition and language identification. Following the concept of SUPERB, ML-SUPERB utilizes frozen SSL features and employs a simple framework for multilingual tasks by learning a shallow downstream model. Similar to the SUPERB benchmark, we find speech SSL models can significantly improve performance compared to FBANK features. Furthermore, we find that multilingual models do not always perform better than their monolingual counterparts. We will release ML-SUPERB as a challenge with organized datasets and reproducible training scripts for future multilingual representation research. |
| title | ML-SUPERB: Multilingual Speech Universal PERformance Benchmark |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2305.10615 |