Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhawan, Kunal, Koluguri, Nithin Rao, Jukić, Ante, Langman, Ryan, Balam, Jagadeesh, Ginsburg, Boris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910619005878272
author Dhawan, Kunal
Koluguri, Nithin Rao
Jukić, Ante
Langman, Ryan
Balam, Jagadeesh
Ginsburg, Boris
author_facet Dhawan, Kunal
Koluguri, Nithin Rao
Jukić, Ante
Langman, Ryan
Balam, Jagadeesh
Ginsburg, Boris
contents Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech recognition (ASR), translation, speaker verification, and joint speech-text foundational models. In this work, we present a comprehensive analysis on building ASR systems with discrete codes. We investigate different methods for codec training such as quantization schemes and time-domain vs spectral feature encodings. We further explore ASR training techniques aimed at enhancing performance, training efficiency, and noise robustness. Drawing upon our findings, we introduce a codec ASR pipeline that outperforms Encodec at similar bit-rate. Remarkably, it also surpasses the state-of-the-art results achieved by strong self-supervised models on the 143 languages ML-SUPERB benchmark despite being smaller in size and pretrained on significantly less data.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03495
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
Dhawan, Kunal
Koluguri, Nithin Rao
Jukić, Ante
Langman, Ryan
Balam, Jagadeesh
Ginsburg, Boris
Audio and Speech Processing
Computation and Language
Machine Learning
Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech recognition (ASR), translation, speaker verification, and joint speech-text foundational models. In this work, we present a comprehensive analysis on building ASR systems with discrete codes. We investigate different methods for codec training such as quantization schemes and time-domain vs spectral feature encodings. We further explore ASR training techniques aimed at enhancing performance, training efficiency, and noise robustness. Drawing upon our findings, we introduce a codec ASR pipeline that outperforms Encodec at similar bit-rate. Remarkably, it also surpasses the state-of-the-art results achieved by strong self-supervised models on the 143 languages ML-SUPERB benchmark despite being smaller in size and pretrained on significantly less data.
title Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
topic Audio and Speech Processing
Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.03495