Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Puvvada, Krishna C., Żelasko, Piotr, Huang, He, Hrinchuk, Oleksii, Koluguri, Nithin Rao, Dhawan, Kunal, Majumdar, Somshubra, Rastorgueva, Elena, Chen, Zhehuai, Lavrukhin, Vitaly, Balam, Jagadeesh, Ginsburg, Boris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913407092916224
author Puvvada, Krishna C.
Żelasko, Piotr
Huang, He
Hrinchuk, Oleksii
Koluguri, Nithin Rao
Dhawan, Kunal
Majumdar, Somshubra
Rastorgueva, Elena
Chen, Zhehuai
Lavrukhin, Vitaly
Balam, Jagadeesh
Ginsburg, Boris
author_facet Puvvada, Krishna C.
Żelasko, Piotr
Huang, He
Hrinchuk, Oleksii
Koluguri, Nithin Rao
Dhawan, Kunal
Majumdar, Somshubra
Rastorgueva, Elena
Chen, Zhehuai
Lavrukhin, Vitaly
Balam, Jagadeesh
Ginsburg, Boris
contents Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data. We argue that state-of-the art accuracy can be reached without relying on web-scale data. Canary - multilingual ASR and speech translation model, outperforms current state-of-the-art models - Whisper, OWSM, and Seamless-M4T on English, French, Spanish, and German languages, while being trained on an order of magnitude less data than these models. Three key factors enables such data-efficient model: (1) a FastConformer-based attention encoder-decoder architecture (2) training on synthetic data generated with machine translation and (3) advanced training techniques: data-balancing, dynamic data blending, dynamic bucketing and noise-robust fine-tuning. The model, weights, and training code will be open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19674
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
Puvvada, Krishna C.
Żelasko, Piotr
Huang, He
Hrinchuk, Oleksii
Koluguri, Nithin Rao
Dhawan, Kunal
Majumdar, Somshubra
Rastorgueva, Elena
Chen, Zhehuai
Lavrukhin, Vitaly
Balam, Jagadeesh
Ginsburg, Boris
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data. We argue that state-of-the art accuracy can be reached without relying on web-scale data. Canary - multilingual ASR and speech translation model, outperforms current state-of-the-art models - Whisper, OWSM, and Seamless-M4T on English, French, Spanish, and German languages, while being trained on an order of magnitude less data than these models. Three key factors enables such data-efficient model: (1) a FastConformer-based attention encoder-decoder architecture (2) training on synthetic data generated with machine translation and (3) advanced training techniques: data-balancing, dynamic data blending, dynamic bucketing and noise-robust fine-tuning. The model, weights, and training code will be open-sourced.
title Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.19674