Dealing with the Hard Facts of Low-Resource African NLP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Diarra, Yacouba, Coulibaly, Nouhoum Souleymane, Kamaté, Panga Azazia, Tall, Madani Amadou, Koné, Emmanuel Élisé, Dembélé, Aymane, Leventhal, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915634716082176
author Diarra, Yacouba
Coulibaly, Nouhoum Souleymane
Kamaté, Panga Azazia
Tall, Madani Amadou
Koné, Emmanuel Élisé
Dembélé, Aymane
Leventhal, Michael
author_facet Diarra, Yacouba
Coulibaly, Nouhoum Souleymane
Kamaté, Panga Azazia
Tall, Madani Amadou
Koné, Emmanuel Élisé
Dembélé, Aymane
Leventhal, Michael
contents Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of spontaneous speech in Bambara, a low-resource West African language; the semi-automated annotation of that dataset with transcriptions; the creation of several monolingual ultra-compact and small models using the dataset; and the automatic and human evaluation of their output. We offer practical suggestions for data collection protocols, annotation, and model design, as well as evidence for the importance of performing human evaluation. In addition to the main dataset, multiple evaluation datasets, models, and code are made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dealing with the Hard Facts of Low-Resource African NLP
Diarra, Yacouba
Coulibaly, Nouhoum Souleymane
Kamaté, Panga Azazia
Tall, Madani Amadou
Koné, Emmanuel Élisé
Dembélé, Aymane
Leventhal, Michael
Computation and Language
Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of spontaneous speech in Bambara, a low-resource West African language; the semi-automated annotation of that dataset with transcriptions; the creation of several monolingual ultra-compact and small models using the dataset; and the automatic and human evaluation of their output. We offer practical suggestions for data collection protocols, annotation, and model design, as well as evidence for the importance of performing human evaluation. In addition to the main dataset, multiple evaluation datasets, models, and code are made publicly available.
title Dealing with the Hard Facts of Low-Resource African NLP
topic Computation and Language
url https://arxiv.org/abs/2511.18557