FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
Fuente:
arXiv
Saved in:
| Main Authors: | Papi, Sara, Gaido, Marco, Bentivogli, Luisa, Brutti, Alessio, Cettolo, Mauro, Gretter, Roberto, Matassoni, Marco, Nabih, Mohamed, Negri, Matteo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
by: Gaido, Marco, et al.
Published: (2025)
by: Gaido, Marco, et al.
Published: (2025)
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
by: Gaido, Marco, et al.
Published: (2025)
by: Gaido, Marco, et al.
Published: (2025)
Cross-Attention is Half Explanation in Speech-to-Text Models
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
by: Fucci, Dennis, et al.
Published: (2024)
by: Fucci, Dennis, et al.
Published: (2024)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
by: Cettolo, Mauro, et al.
Published: (2025)
by: Cettolo, Mauro, et al.
Published: (2025)
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena
by: Gaido, Marco, et al.
Published: (2024)
by: Gaido, Marco, et al.
Published: (2024)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
by: Lam, Tsz Kin, et al.
Published: (2025)
by: Lam, Tsz Kin, et al.
Published: (2025)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
by: Papi, Sara, et al.
Published: (2026)
by: Papi, Sara, et al.
Published: (2026)
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
by: Fucci, Dennis, et al.
Published: (2025)
by: Fucci, Dennis, et al.
Published: (2025)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2023)
by: Papi, Sara, et al.
Published: (2023)
MLMA: Towards Multilingual ASR With Mamba-based Architectures
by: Ali, Mohamed Nabih, et al.
Published: (2025)
by: Ali, Mohamed Nabih, et al.
Published: (2025)
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
The Eloquence team submission for task 1 of MLC-SLM challenge
by: Concina, Lorenzo, et al.
Published: (2025)
by: Concina, Lorenzo, et al.
Published: (2025)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
by: Djanibekov, Amirbek, et al.
Published: (2026)
by: Djanibekov, Amirbek, et al.
Published: (2026)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
by: Fong, Seraphina, et al.
Published: (2025)
by: Fong, Seraphina, et al.
Published: (2025)
Input Conditioned Layer Dropping in Speech Foundation Models
by: Hannan, Abdul, et al.
Published: (2025)
by: Hannan, Abdul, et al.
Published: (2025)
When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP
by: Papi, Sara, et al.
Published: (2023)
by: Papi, Sara, et al.
Published: (2023)
The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
by: Conti, Lina, et al.
Published: (2025)
by: Conti, Lina, et al.
Published: (2025)
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
by: Conti, Lina, et al.
Published: (2025)
by: Conti, Lina, et al.
Published: (2025)
Different Speech Translation Models Encode and Translate Speaker Gender Differently
by: Fucci, Dennis, et al.
Published: (2025)
by: Fucci, Dennis, et al.
Published: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
by: Wright, George August, et al.
Published: (2023)
by: Wright, George August, et al.
Published: (2023)
Federating Dynamic Models using Early-Exit Architectures for Automatic Speech Recognition on Heterogeneous Clients
by: Ali, Mohamed Nabih, et al.
Published: (2024)
by: Ali, Mohamed Nabih, et al.
Published: (2024)
NUTSHELL: A Dataset for Abstract Generation from Scientific Talks
by: Züfle, Maike, et al.
Published: (2025)
by: Züfle, Maike, et al.
Published: (2025)
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study
by: Savoldi, Beatrice, et al.
Published: (2024)
by: Savoldi, Beatrice, et al.
Published: (2024)
Distillation-based Layer Dropping (DLD): Effective End-to-end Framework for Dynamic Speech Networks
by: Hannan, Abdul, et al.
Published: (2026)
by: Hannan, Abdul, et al.
Published: (2026)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
by: Morrone, Giovanni, et al.
Published: (2023)
by: Morrone, Giovanni, et al.
Published: (2023)
Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
by: Piergentili, Andrea, et al.
Published: (2025)
by: Piergentili, Andrea, et al.
Published: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities
by: Qian, Xinyuan, et al.
Published: (2026)
by: Qian, Xinyuan, et al.
Published: (2026)
OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
by: Liao, Yujie, et al.
Published: (2026)
by: Liao, Yujie, et al.
Published: (2026)
Hallucination Benchmark for Speech Foundation Models
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
by: Verdini, Francesco, et al.
Published: (2024)
by: Verdini, Francesco, et al.
Published: (2024)
Similar Items
-
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
by: Gaido, Marco, et al.
Published: (2024) -
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
by: Gaido, Marco, et al.
Published: (2025) -
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2024) -
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
by: Papi, Sara, et al.
Published: (2024) -
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
by: Gaido, Marco, et al.
Published: (2025)