Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Chien-yu, Lu, Ke-Han, Wang, Shih-Heng, Hsiao, Chi-Yuan, Kuan, Chun-Yi, Wu, Haibin, Arora, Siddhant, Chang, Kai-Wei, Shi, Jiatong, Peng, Yifan, Sharma, Roshan, Watanabe, Shinji, Ramakrishnan, Bhiksha, Shehata, Shady, Lee, Hung-yi
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:https://arxiv.org/abs/2309.09510
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910378258071552
author Huang, Chien-yu
Lu, Ke-Han
Wang, Shih-Heng
Hsiao, Chi-Yuan
Kuan, Chun-Yi
Wu, Haibin
Arora, Siddhant
Chang, Kai-Wei
Shi, Jiatong
Peng, Yifan
Sharma, Roshan
Watanabe, Shinji
Ramakrishnan, Bhiksha
Shehata, Shady
Lee, Hung-yi
author_facet Huang, Chien-yu
Lu, Ke-Han
Wang, Shih-Heng
Hsiao, Chi-Yuan
Kuan, Chun-Yi
Wu, Haibin
Arora, Siddhant
Chang, Kai-Wei
Shi, Jiatong
Peng, Yifan
Sharma, Roshan
Watanabe, Shinji
Ramakrishnan, Bhiksha
Shehata, Shady
Lee, Hung-yi
contents Text language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair comparison across different approaches. Thus, we present Dynamic-SUPERB, a benchmark designed for building universal speech models capable of leveraging instruction tuning to perform multiple tasks in a zero-shot fashion. To achieve comprehensive coverage of diverse speech tasks and harness instruction tuning, we invite the community to collaborate and contribute, facilitating the dynamic growth of the benchmark. To initiate, Dynamic-SUPERB features 55 evaluation instances by combining 33 tasks and 22 datasets. This spans a broad spectrum of dimensions, providing a comprehensive platform for evaluation. Additionally, we propose several approaches to establish benchmark baselines. These include the utilization of speech models, text language models, and the multimodal encoder. Evaluation results indicate that while these baselines perform reasonably on seen tasks, they struggle with unseen ones. We release all materials to the public and welcome researchers to collaborate on the project, advancing technologies in the field together.
format Preprint
id arxiv_https___arxiv_org_abs_2309_09510
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
Huang, Chien-yu
Lu, Ke-Han
Wang, Shih-Heng
Hsiao, Chi-Yuan
Kuan, Chun-Yi
Wu, Haibin
Arora, Siddhant
Chang, Kai-Wei
Shi, Jiatong
Peng, Yifan
Sharma, Roshan
Watanabe, Shinji
Ramakrishnan, Bhiksha
Shehata, Shady
Lee, Hung-yi
Audio and Speech Processing
Machine Learning
Sound
Text language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair comparison across different approaches. Thus, we present Dynamic-SUPERB, a benchmark designed for building universal speech models capable of leveraging instruction tuning to perform multiple tasks in a zero-shot fashion. To achieve comprehensive coverage of diverse speech tasks and harness instruction tuning, we invite the community to collaborate and contribute, facilitating the dynamic growth of the benchmark. To initiate, Dynamic-SUPERB features 55 evaluation instances by combining 33 tasks and 22 datasets. This spans a broad spectrum of dimensions, providing a comprehensive platform for evaluation. Additionally, we propose several approaches to establish benchmark baselines. These include the utilization of speech models, text language models, and the multimodal encoder. Evaluation results indicate that while these baselines perform reasonably on seen tasks, they struggle with unseen ones. We release all materials to the public and welcome researchers to collaborate on the project, advancing technologies in the field together.
title Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2309.09510