MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gaido, Marco, Papi, Sara, Bentivogli, Luisa, Brutti, Alessio, Cettolo, Mauro, Gretter, Roberto, Matassoni, Marco, Nabih, Mohamed, Negri, Matteo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
von: Gaido, Marco, et al.
Veröffentlicht: (2025)
von: Gaido, Marco, et al.
Veröffentlicht: (2025)
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
von: Gaido, Marco, et al.
Veröffentlicht: (2025)
von: Gaido, Marco, et al.
Veröffentlicht: (2025)
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
von: Fucci, Dennis, et al.
Veröffentlicht: (2024)
von: Fucci, Dennis, et al.
Veröffentlicht: (2024)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
von: Cettolo, Mauro, et al.
Veröffentlicht: (2025)
von: Cettolo, Mauro, et al.
Veröffentlicht: (2025)
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
von: Fong, Seraphina, et al.
Veröffentlicht: (2025)
von: Fong, Seraphina, et al.
Veröffentlicht: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2025)
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2025)
Cross-Attention is Half Explanation in Speech-to-Text Models
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
von: Fucci, Dennis, et al.
Veröffentlicht: (2025)
von: Fucci, Dennis, et al.
Veröffentlicht: (2025)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
von: Papi, Sara, et al.
Veröffentlicht: (2023)
von: Papi, Sara, et al.
Veröffentlicht: (2023)
The Eloquence team submission for task 1 of MLC-SLM challenge
von: Concina, Lorenzo, et al.
Veröffentlicht: (2025)
von: Concina, Lorenzo, et al.
Veröffentlicht: (2025)
Input Conditioned Layer Dropping in Speech Foundation Models
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
Granary: Speech Recognition and Translation Dataset in 25 European Languages
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2025)
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2025)
The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
von: Conti, Lina, et al.
Veröffentlicht: (2025)
von: Conti, Lina, et al.
Veröffentlicht: (2025)
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
von: Conti, Lina, et al.
Veröffentlicht: (2025)
von: Conti, Lina, et al.
Veröffentlicht: (2025)
Different Speech Translation Models Encode and Translate Speaker Gender Differently
von: Fucci, Dennis, et al.
Veröffentlicht: (2025)
von: Fucci, Dennis, et al.
Veröffentlicht: (2025)
Federating Dynamic Models using Early-Exit Architectures for Automatic Speech Recognition on Heterogeneous Clients
von: Ali, Mohamed Nabih, et al.
Veröffentlicht: (2024)
von: Ali, Mohamed Nabih, et al.
Veröffentlicht: (2024)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
von: Papi, Sara, et al.
Veröffentlicht: (2026)
von: Papi, Sara, et al.
Veröffentlicht: (2026)
Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
von: Chen, Yushen, et al.
Veröffentlicht: (2026)
von: Chen, Yushen, et al.
Veröffentlicht: (2026)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
von: Morrone, Giovanni, et al.
Veröffentlicht: (2023)
von: Morrone, Giovanni, et al.
Veröffentlicht: (2023)
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
Open-Source System for Multilingual Translation and Cloned Speech Synthesis
von: Cámara, Mateo, et al.
Veröffentlicht: (2025)
von: Cámara, Mateo, et al.
Veröffentlicht: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP
von: Papi, Sara, et al.
Veröffentlicht: (2023)
von: Papi, Sara, et al.
Veröffentlicht: (2023)
Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
von: Chhibber, Manasi, et al.
Veröffentlicht: (2025)
von: Chhibber, Manasi, et al.
Veröffentlicht: (2025)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
MAD Speech: Measures of Acoustic Diversity of Speech
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
Open Source State-Of-the-Art Solution for Romanian Speech Recognition
von: Pirlogeanu, Gabriel, et al.
Veröffentlicht: (2025)
von: Pirlogeanu, Gabriel, et al.
Veröffentlicht: (2025)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
von: Papi, Sara, et al.
Veröffentlicht: (2025) -
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
von: Gaido, Marco, et al.
Veröffentlicht: (2025) -
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
von: Gaido, Marco, et al.
Veröffentlicht: (2025) -
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
von: Papi, Sara, et al.
Veröffentlicht: (2024) -
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
von: Fucci, Dennis, et al.
Veröffentlicht: (2024)