Dealing with the Hard Facts of Low-Resource African NLP
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915634716082176 |
|---|---|
| author | Diarra, Yacouba Coulibaly, Nouhoum Souleymane Kamaté, Panga Azazia Tall, Madani Amadou Koné, Emmanuel Élisé Dembélé, Aymane Leventhal, Michael |
| author_facet | Diarra, Yacouba Coulibaly, Nouhoum Souleymane Kamaté, Panga Azazia Tall, Madani Amadou Koné, Emmanuel Élisé Dembélé, Aymane Leventhal, Michael |
| contents | Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of spontaneous speech in Bambara, a low-resource West African language; the semi-automated annotation of that dataset with transcriptions; the creation of several monolingual ultra-compact and small models using the dataset; and the automatic and human evaluation of their output. We offer practical suggestions for data collection protocols, annotation, and model design, as well as evidence for the importance of performing human evaluation. In addition to the main dataset, multiple evaluation datasets, models, and code are made publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_18557 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Dealing with the Hard Facts of Low-Resource African NLP Diarra, Yacouba Coulibaly, Nouhoum Souleymane Kamaté, Panga Azazia Tall, Madani Amadou Koné, Emmanuel Élisé Dembélé, Aymane Leventhal, Michael Computation and Language Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of spontaneous speech in Bambara, a low-resource West African language; the semi-automated annotation of that dataset with transcriptions; the creation of several monolingual ultra-compact and small models using the dataset; and the automatic and human evaluation of their output. We offer practical suggestions for data collection protocols, annotation, and model design, as well as evidence for the importance of performing human evaluation. In addition to the main dataset, multiple evaluation datasets, models, and code are made publicly available. |
| title | Dealing with the Hard Facts of Low-Resource African NLP |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2511.18557 |