Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910619005878272 |
|---|---|
| author | Dhawan, Kunal Koluguri, Nithin Rao Jukić, Ante Langman, Ryan Balam, Jagadeesh Ginsburg, Boris |
| author_facet | Dhawan, Kunal Koluguri, Nithin Rao Jukić, Ante Langman, Ryan Balam, Jagadeesh Ginsburg, Boris |
| contents | Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech recognition (ASR), translation, speaker verification, and joint speech-text foundational models. In this work, we present a comprehensive analysis on building ASR systems with discrete codes. We investigate different methods for codec training such as quantization schemes and time-domain vs spectral feature encodings. We further explore ASR training techniques aimed at enhancing performance, training efficiency, and noise robustness. Drawing upon our findings, we introduce a codec ASR pipeline that outperforms Encodec at similar bit-rate. Remarkably, it also surpasses the state-of-the-art results achieved by strong self-supervised models on the 143 languages ML-SUPERB benchmark despite being smaller in size and pretrained on significantly less data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_03495 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations Dhawan, Kunal Koluguri, Nithin Rao Jukić, Ante Langman, Ryan Balam, Jagadeesh Ginsburg, Boris Audio and Speech Processing Computation and Language Machine Learning Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech recognition (ASR), translation, speaker verification, and joint speech-text foundational models. In this work, we present a comprehensive analysis on building ASR systems with discrete codes. We investigate different methods for codec training such as quantization schemes and time-domain vs spectral feature encodings. We further explore ASR training techniques aimed at enhancing performance, training efficiency, and noise robustness. Drawing upon our findings, we introduce a codec ASR pipeline that outperforms Encodec at similar bit-rate. Remarkably, it also surpasses the state-of-the-art results achieved by strong self-supervised models on the 143 languages ML-SUPERB benchmark despite being smaller in size and pretrained on significantly less data. |
| title | Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations |
| topic | Audio and Speech Processing Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2407.03495 |