Voice of a Continent: Mapping Africa's Speech Technology Frontier

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Elmadany, AbdelRahim, Kwon, Sang Yun, Toyin, Hawau Olamide, Inciarte, Alcides Alcoba, Aldarmaki, Hanan, Abdul-Mageed, Muhammad
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913925820317696
author Elmadany, AbdelRahim
Kwon, Sang Yun
Toyin, Hawau Olamide
Inciarte, Alcides Alcoba
Aldarmaki, Hanan
Abdul-Mageed, Muhammad
author_facet Elmadany, AbdelRahim
Kwon, Sang Yun
Toyin, Hawau Olamide
Inciarte, Alcides Alcoba
Aldarmaki, Hanan
Abdul-Mageed, Muhammad
contents Africa's rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent's speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench for downstream African speech tasks. Using SimbaBench, we introduce the Simba family of models, achieving state-of-the-art performance across multiple African languages and speech tasks. Our benchmark analysis reveals critical patterns in resource availability, while our model evaluation demonstrates how dataset quality, domain diversity, and language family relationships influence performance across languages. Our work highlights the need for expanded speech technology resources that better reflect Africa's linguistic diversity and provides a solid foundation for future research and development efforts toward more inclusive speech technologies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Voice of a Continent: Mapping Africa's Speech Technology Frontier
Elmadany, AbdelRahim
Kwon, Sang Yun
Toyin, Hawau Olamide
Inciarte, Alcides Alcoba
Aldarmaki, Hanan
Abdul-Mageed, Muhammad
Computation and Language
Africa's rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent's speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench for downstream African speech tasks. Using SimbaBench, we introduce the Simba family of models, achieving state-of-the-art performance across multiple African languages and speech tasks. Our benchmark analysis reveals critical patterns in resource availability, while our model evaluation demonstrates how dataset quality, domain diversity, and language family relationships influence performance across languages. Our work highlights the need for expanded speech technology resources that better reflect Africa's linguistic diversity and provides a solid foundation for future research and development efforts toward more inclusive speech technologies.
title Voice of a Continent: Mapping Africa's Speech Technology Frontier
topic Computation and Language
url https://arxiv.org/abs/2505.18436