k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912286632837120 |
|---|---|
| author | Yang, Yifan Zhuo, Jianheng Jin, Zengrui Ma, Ziyang Yang, Xiaoyu Yao, Zengwei Guo, Liyong Kang, Wei Kuang, Fangjun Lin, Long Povey, Daniel Chen, Xie |
| author_facet | Yang, Yifan Zhuo, Jianheng Jin, Zengrui Ma, Ziyang Yang, Xiaoyu Yao, Zengwei Guo, Liyong Kang, Wei Kuang, Fangjun Lin, Long Povey, Daniel Chen, Xie |
| contents | Self-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipformer, which excel in automatic speech recognition (ASR), remain unexplored in SSL. Concurrently, inefficiencies in data processing within existing SSL training frameworks, such as fairseq, pose challenges in managing the growing volumes of training data. To address these issues, we propose k2SSL, an open-source framework that offers faster, more memory-efficient, and better-performing self-supervised speech representation learning, focusing on downstream ASR tasks. The optimized HuBERT and proposed Zipformer-based SSL systems exhibit substantial reductions in both training time and memory usage during SSL training. Experiments on LibriSpeech demonstrate that Zipformer Base significantly outperforms HuBERT and WavLM, achieving up to a 34.8% relative WER reduction compared to HuBERT Base after fine-tuning, along with a 3.5x pre-training speedup in GPU hours. When scaled to 60k hours of LibriLight data, Zipformer Large exhibits remarkable efficiency, matching HuBERT Large's performance while requiring only 5/8 pre-training steps. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_17100 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning Yang, Yifan Zhuo, Jianheng Jin, Zengrui Ma, Ziyang Yang, Xiaoyu Yao, Zengwei Guo, Liyong Kang, Wei Kuang, Fangjun Lin, Long Povey, Daniel Chen, Xie Audio and Speech Processing Self-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipformer, which excel in automatic speech recognition (ASR), remain unexplored in SSL. Concurrently, inefficiencies in data processing within existing SSL training frameworks, such as fairseq, pose challenges in managing the growing volumes of training data. To address these issues, we propose k2SSL, an open-source framework that offers faster, more memory-efficient, and better-performing self-supervised speech representation learning, focusing on downstream ASR tasks. The optimized HuBERT and proposed Zipformer-based SSL systems exhibit substantial reductions in both training time and memory usage during SSL training. Experiments on LibriSpeech demonstrate that Zipformer Base significantly outperforms HuBERT and WavLM, achieving up to a 34.8% relative WER reduction compared to HuBERT Base after fine-tuning, along with a 3.5x pre-training speedup in GPU hours. When scaled to 60k hours of LibriLight data, Zipformer Large exhibits remarkable efficiency, matching HuBERT Large's performance while requiring only 5/8 pre-training steps. |
| title | k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2411.17100 |