FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Papi, Sara, Gaido, Marco, Bentivogli, Luisa, Brutti, Alessio, Cettolo, Mauro, Gretter, Roberto, Matassoni, Marco, Nabih, Mohamed, Negri, Matteo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916770242101248
author Papi, Sara
Gaido, Marco
Bentivogli, Luisa
Brutti, Alessio
Cettolo, Mauro
Gretter, Roberto
Matassoni, Marco
Nabih, Mohamed
Negri, Matteo
author_facet Papi, Sara
Gaido, Marco
Bentivogli, Luisa
Brutti, Alessio
Cettolo, Mauro
Gretter, Roberto
Matassoni, Marco
Nabih, Mohamed
Negri, Matteo
contents The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inaccessible training data and code--poses major reproducibility and fair evaluation challenges. While other domains have made substantial progress toward open science by developing fully transparent models trained on open-source (OS) code and data, similar efforts in speech remain limited. To fill this gap, we introduce FAMA, the first family of open science SFMs for English and Italian, trained on 150k+ hours of OS speech data. Moreover, we present a new dataset containing 16k hours of cleaned and pseudo-labeled speech for both languages. Results show that FAMA achieves competitive performance compared to existing SFMs while being up to 8 times faster. All artifacts, including code, datasets, and models, are released under OS-compliant licenses, promoting openness in speech technology research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22759
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
Papi, Sara
Gaido, Marco
Bentivogli, Luisa
Brutti, Alessio
Cettolo, Mauro
Gretter, Roberto
Matassoni, Marco
Nabih, Mohamed
Negri, Matteo
Computation and Language
Artificial Intelligence
Sound
The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inaccessible training data and code--poses major reproducibility and fair evaluation challenges. While other domains have made substantial progress toward open science by developing fully transparent models trained on open-source (OS) code and data, similar efforts in speech remain limited. To fill this gap, we introduce FAMA, the first family of open science SFMs for English and Italian, trained on 150k+ hours of OS speech data. Moreover, we present a new dataset containing 16k hours of cleaned and pseudo-labeled speech for both languages. Results show that FAMA achieves competitive performance compared to existing SFMs while being up to 8 times faster. All artifacts, including code, datasets, and models, are released under OS-compliant licenses, promoting openness in speech technology research.
title FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
topic Computation and Language
Artificial Intelligence
Sound
url https://arxiv.org/abs/2505.22759