Saved in:
Bibliographic Details
Main Authors: Yang, Shu-wen, Chang, Heng-Jui, Huang, Zili, Liu, Andy T., Lai, Cheng-I, Wu, Haibin, Shi, Jiatong, Chang, Xuankai, Tsai, Hsiang-Sheng, Huang, Wen-Chin, Feng, Tzu-hsun, Chi, Po-Han, Lin, Yist Y., Chuang, Yung-Sung, Huang, Tzu-Hsien, Tseng, Wei-Cheng, Lakhotia, Kushal, Li, Shang-Wen, Mohamed, Abdelrahman, Watanabe, Shinji, Lee, Hung-yi
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.09385
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916265602318336
author Yang, Shu-wen
Chang, Heng-Jui
Huang, Zili
Liu, Andy T.
Lai, Cheng-I
Wu, Haibin
Shi, Jiatong
Chang, Xuankai
Tsai, Hsiang-Sheng
Huang, Wen-Chin
Feng, Tzu-hsun
Chi, Po-Han
Lin, Yist Y.
Chuang, Yung-Sung
Huang, Tzu-Hsien
Tseng, Wei-Cheng
Lakhotia, Kushal
Li, Shang-Wen
Mohamed, Abdelrahman
Watanabe, Shinji
Lee, Hung-yi
author_facet Yang, Shu-wen
Chang, Heng-Jui
Huang, Zili
Liu, Andy T.
Lai, Cheng-I
Wu, Haibin
Shi, Jiatong
Chang, Xuankai
Tsai, Hsiang-Sheng
Huang, Wen-Chin
Feng, Tzu-hsun
Chi, Po-Han
Lin, Yist Y.
Chuang, Yung-Sung
Huang, Tzu-Hsien
Tseng, Wei-Cheng
Lakhotia, Kushal
Li, Shang-Wen
Mohamed, Abdelrahman
Watanabe, Shinji
Lee, Hung-yi
contents The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. In this work, we establish the Speech processing Universal PERformance Benchmark (SUPERB) to study the effectiveness of the paradigm for speech. We propose a unified multi-tasking framework to address speech processing tasks in SUPERB using a frozen foundation model followed by task-specialized, lightweight prediction heads. Combining our results with community submissions, we verify that the foundation model paradigm is promising for speech, and our multi-tasking framework is simple yet effective, as the best-performing foundation model shows competitive generalizability across most SUPERB tasks. For reproducibility and extensibility, we have developed a long-term maintained platform that enables deterministic benchmarking, allows for result sharing via an online leaderboard, and promotes collaboration through a community-driven benchmark database to support new development cycles. Finally, we conduct a series of analyses to offer an in-depth understanding of SUPERB and speech foundation models, including information flows across tasks inside the models, the correctness of the weighted-sum benchmarking protocol and the statistical significance and robustness of the benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09385
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Large-Scale Evaluation of Speech Foundation Models
Yang, Shu-wen
Chang, Heng-Jui
Huang, Zili
Liu, Andy T.
Lai, Cheng-I
Wu, Haibin
Shi, Jiatong
Chang, Xuankai
Tsai, Hsiang-Sheng
Huang, Wen-Chin
Feng, Tzu-hsun
Chi, Po-Han
Lin, Yist Y.
Chuang, Yung-Sung
Huang, Tzu-Hsien
Tseng, Wei-Cheng
Lakhotia, Kushal
Li, Shang-Wen
Mohamed, Abdelrahman
Watanabe, Shinji
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Signal Processing
The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. In this work, we establish the Speech processing Universal PERformance Benchmark (SUPERB) to study the effectiveness of the paradigm for speech. We propose a unified multi-tasking framework to address speech processing tasks in SUPERB using a frozen foundation model followed by task-specialized, lightweight prediction heads. Combining our results with community submissions, we verify that the foundation model paradigm is promising for speech, and our multi-tasking framework is simple yet effective, as the best-performing foundation model shows competitive generalizability across most SUPERB tasks. For reproducibility and extensibility, we have developed a long-term maintained platform that enables deterministic benchmarking, allows for result sharing via an online leaderboard, and promotes collaboration through a community-driven benchmark database to support new development cycles. Finally, we conduct a series of analyses to offer an in-depth understanding of SUPERB and speech foundation models, including information flows across tasks inside the models, the correctness of the weighted-sum benchmarking protocol and the statistical significance and robustness of the benchmark.
title A Large-Scale Evaluation of Speech Foundation Models
topic Audio and Speech Processing
Computation and Language
Signal Processing
url https://arxiv.org/abs/2404.09385