From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Tianduo, Xu, Lu, Lu, Wei, Cheng, Shanbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916752782262272
author Wang, Tianduo
Xu, Lu
Lu, Wei
Cheng, Shanbo
author_facet Wang, Tianduo
Xu, Lu
Lu, Wei
Cheng, Shanbo
contents Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into synthetic speech via off-the-shelf text-to-speech (TTS) models. We demonstrate that just tens of hours of real transcribed speech can effectively train TTS models to generate synthetic speech at hundreds of times the original volume while maintaining high quality. To evaluate synthetic speech quality, we develop an intelligibility-based assessment framework and establish clear thresholds for when synthetic data benefits ASR training. Using Speech Back-Translation, we generate more than 500,000 hours of synthetic speech in ten languages and continue pre-training Whisper-large-v3, achieving average transcription error reductions of over 30\%. These results highlight the scalability and effectiveness of Speech Back-Translation for enhancing multilingual ASR systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16972
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Wang, Tianduo
Xu, Lu
Lu, Wei
Cheng, Shanbo
Computation and Language
Sound
Audio and Speech Processing
Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into synthetic speech via off-the-shelf text-to-speech (TTS) models. We demonstrate that just tens of hours of real transcribed speech can effectively train TTS models to generate synthetic speech at hundreds of times the original volume while maintaining high quality. To evaluate synthetic speech quality, we develop an intelligibility-based assessment framework and establish clear thresholds for when synthetic data benefits ASR training. Using Speech Back-Translation, we generate more than 500,000 hours of synthetic speech in ten languages and continue pre-training Whisper-large-v3, achieving average transcription error reductions of over 30\%. These results highlight the scalability and effectiveness of Speech Back-Translation for enhancing multilingual ASR systems.
title From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.16972