Synthetic Voice Data for Automatic Speech Recognition in African Languages
Fuente:
arXiv
Saved in:
| Main Authors: | DeRenzi, Brian, Dixon, Anna, Farhi, Mohamed Aymane, Resch, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
by: Lajčinová, Bibiána, et al.
Published: (2024)
by: Lajčinová, Bibiána, et al.
Published: (2024)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
by: Klöser, Lars, et al.
Published: (2024)
by: Klöser, Lars, et al.
Published: (2024)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
by: Chowdhury, MD. Sagor, et al.
Published: (2026)
by: Chowdhury, MD. Sagor, et al.
Published: (2026)
Automatic End-to-End Data Integration using Large Language Models
by: Steiner, Aaron, et al.
Published: (2026)
by: Steiner, Aaron, et al.
Published: (2026)
Curating Grounded Synthetic Data with Global Perspectives for Equitable AI
by: Törnquist, Elin, et al.
Published: (2024)
by: Törnquist, Elin, et al.
Published: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
by: Kuhn, Korbinian, et al.
Published: (2024)
by: Kuhn, Korbinian, et al.
Published: (2024)
Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE
by: Dahl, Christian Møller, et al.
Published: (2024)
by: Dahl, Christian Møller, et al.
Published: (2024)
Budget-Xfer: Budget-Constrained Source Language Selection for Cross-Lingual Transfer to African Languages
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
BlasBench: An Open Benchmark for Irish Speech Recognition
by: Raj, Jyoutir, et al.
Published: (2026)
by: Raj, Jyoutir, et al.
Published: (2026)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
by: Kang, Migyeong, et al.
Published: (2026)
by: Kang, Migyeong, et al.
Published: (2026)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
by: Sauter, Adrian, et al.
Published: (2025)
by: Sauter, Adrian, et al.
Published: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
by: Ahn, Taekyung, et al.
Published: (2024)
by: Ahn, Taekyung, et al.
Published: (2024)
SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models
by: Rubinstein, Beny, et al.
Published: (2026)
by: Rubinstein, Beny, et al.
Published: (2026)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
Identifying Fairness Issues in Automatically Generated Testing Content
by: Stowe, Kevin, et al.
Published: (2024)
by: Stowe, Kevin, et al.
Published: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
by: Ge, Danying, et al.
Published: (2025)
by: Ge, Danying, et al.
Published: (2025)
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
by: Signoroni, Edoardo, et al.
Published: (2026)
by: Signoroni, Edoardo, et al.
Published: (2026)
Automatic Generation of Conversational Interfaces for Tabular Data Analysis
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
Socially Responsible Data for Large Multilingual Language Models
by: Smart, Andrew, et al.
Published: (2024)
by: Smart, Andrew, et al.
Published: (2024)
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
by: Hahn, Udo
Published: (2024)
by: Hahn, Udo
Published: (2024)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
by: Yun, Janghyeon, et al.
Published: (2025)
by: Yun, Janghyeon, et al.
Published: (2025)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
Graphemic Normalization of the Perso-Arabic Script
by: Doctor, Raiomond, et al.
Published: (2022)
by: Doctor, Raiomond, et al.
Published: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2026)
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2026)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025)
by: Doghmash, Salam Thabet, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
myNER: Contextualized Burmese Named Entity Recognition with Bidirectional LSTM and fastText Embeddings via Joint Training with POS Tagging
by: Thant, Kaung Lwin, et al.
Published: (2025)
by: Thant, Kaung Lwin, et al.
Published: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025)
by: Zhang, Yizhuo, et al.
Published: (2025)
Similar Items
-
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
by: Lajčinová, Bibiána, et al.
Published: (2024) -
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
by: Klöser, Lars, et al.
Published: (2024) -
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
by: Chowdhury, MD. Sagor, et al.
Published: (2026) -
Automatic End-to-End Data Integration using Large Language Models
by: Steiner, Aaron, et al.
Published: (2026) -
Curating Grounded Synthetic Data with Global Perspectives for Equitable AI
by: Törnquist, Elin, et al.
Published: (2024)