k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Yifan, Zhuo, Jianheng, Jin, Zengrui, Ma, Ziyang, Yang, Xiaoyu, Yao, Zengwei, Guo, Liyong, Kang, Wei, Kuang, Fangjun, Lin, Long, Povey, Daniel, Chen, Xie
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912286632837120
author Yang, Yifan
Zhuo, Jianheng
Jin, Zengrui
Ma, Ziyang
Yang, Xiaoyu
Yao, Zengwei
Guo, Liyong
Kang, Wei
Kuang, Fangjun
Lin, Long
Povey, Daniel
Chen, Xie
author_facet Yang, Yifan
Zhuo, Jianheng
Jin, Zengrui
Ma, Ziyang
Yang, Xiaoyu
Yao, Zengwei
Guo, Liyong
Kang, Wei
Kuang, Fangjun
Lin, Long
Povey, Daniel
Chen, Xie
contents Self-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipformer, which excel in automatic speech recognition (ASR), remain unexplored in SSL. Concurrently, inefficiencies in data processing within existing SSL training frameworks, such as fairseq, pose challenges in managing the growing volumes of training data. To address these issues, we propose k2SSL, an open-source framework that offers faster, more memory-efficient, and better-performing self-supervised speech representation learning, focusing on downstream ASR tasks. The optimized HuBERT and proposed Zipformer-based SSL systems exhibit substantial reductions in both training time and memory usage during SSL training. Experiments on LibriSpeech demonstrate that Zipformer Base significantly outperforms HuBERT and WavLM, achieving up to a 34.8% relative WER reduction compared to HuBERT Base after fine-tuning, along with a 3.5x pre-training speedup in GPU hours. When scaled to 60k hours of LibriLight data, Zipformer Large exhibits remarkable efficiency, matching HuBERT Large's performance while requiring only 5/8 pre-training steps.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17100
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
Yang, Yifan
Zhuo, Jianheng
Jin, Zengrui
Ma, Ziyang
Yang, Xiaoyu
Yao, Zengwei
Guo, Liyong
Kang, Wei
Kuang, Fangjun
Lin, Long
Povey, Daniel
Chen, Xie
Audio and Speech Processing
Self-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipformer, which excel in automatic speech recognition (ASR), remain unexplored in SSL. Concurrently, inefficiencies in data processing within existing SSL training frameworks, such as fairseq, pose challenges in managing the growing volumes of training data. To address these issues, we propose k2SSL, an open-source framework that offers faster, more memory-efficient, and better-performing self-supervised speech representation learning, focusing on downstream ASR tasks. The optimized HuBERT and proposed Zipformer-based SSL systems exhibit substantial reductions in both training time and memory usage during SSL training. Experiments on LibriSpeech demonstrate that Zipformer Base significantly outperforms HuBERT and WavLM, achieving up to a 34.8% relative WER reduction compared to HuBERT Base after fine-tuning, along with a 3.5x pre-training speedup in GPU hours. When scaled to 60k hours of LibriLight data, Zipformer Large exhibits remarkable efficiency, matching HuBERT Large's performance while requiring only 5/8 pre-training steps.
title k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
topic Audio and Speech Processing
url https://arxiv.org/abs/2411.17100