OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, William, Tian, Jinchuan, Peng, Yifan, Yan, Brian, Yang, Chao-Han Huck, Watanabe, Shinji
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909493643706368
author Chen, William
Tian, Jinchuan
Peng, Yifan
Yan, Brian
Yang, Chao-Han Huck
Watanabe, Shinji
author_facet Chen, William
Tian, Jinchuan
Peng, Yifan
Yan, Brian
Yang, Chao-Han Huck
Watanabe, Shinji
contents Neural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In this work, we introduce OWLS, an open-access, reproducible suite of multilingual speech recognition and translation models spanning 0.25B to 18B parameters, with the 18B version being the largest speech model, to the best of our knowledge. OWLS leverages up to 360K hours of public speech data across 150 languages, enabling a systematic investigation into how data, model, and compute scaling each influence performance in multilingual speech tasks. We use OWLS to derive neural scaling laws, showing how final performance can be reliably predicted when scaling. One of our key findings is that scaling enhances performance on low-resource languages/dialects, helping to mitigate bias and improve the accessibility of speech technologies. Finally, we show how OWLS can be used to power new research directions by discovering emergent abilities in large-scale speech models. Model checkpoints will be released on https://huggingface.co/collections/espnet/owls-scaling-laws-for-speech-recognition-and-translation-67ab7f991c194065f057ce8d for future studies.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
Chen, William
Tian, Jinchuan
Peng, Yifan
Yan, Brian
Yang, Chao-Han Huck
Watanabe, Shinji
Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Neural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In this work, we introduce OWLS, an open-access, reproducible suite of multilingual speech recognition and translation models spanning 0.25B to 18B parameters, with the 18B version being the largest speech model, to the best of our knowledge. OWLS leverages up to 360K hours of public speech data across 150 languages, enabling a systematic investigation into how data, model, and compute scaling each influence performance in multilingual speech tasks. We use OWLS to derive neural scaling laws, showing how final performance can be reliably predicted when scaling. One of our key findings is that scaling enhances performance on low-resource languages/dialects, helping to mitigate bias and improve the accessibility of speech technologies. Finally, we show how OWLS can be used to power new research directions by discovering emergent abilities in large-scale speech models. Model checkpoints will be released on https://huggingface.co/collections/espnet/owls-scaling-laws-for-speech-recognition-and-translation-67ab7f991c194065f057ce8d for future studies.
title OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2502.10373