Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Diarra, Yacouba, Kamate, Panga Azazia, Coulibaly, Nouhoum Souleymane, Leventhal, Michael
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912782776008704
author Diarra, Yacouba
Kamate, Panga Azazia
Coulibaly, Nouhoum Souleymane
Leventhal, Michael
author_facet Diarra, Yacouba
Kamate, Panga Azazia
Coulibaly, Nouhoum Souleymane
Leventhal, Michael
contents We present Kunkado, a 160-hour Bambara ASR dataset compiled from Malian radio archives to capture present-day spontaneous speech across a wide range of topics. It includes code-switching, disfluencies, background noise, and overlapping speakers that practical ASR systems encounter in real-world use. We finetuned Parakeet-based models on a 33.47-hour human-reviewed subset and apply pragmatic transcript normalization to reduce variability in number formatting, tags, and code-switching annotations. Evaluated on two real-world test sets, finetuning with Kunkado reduces WER from 44.47\% to 37.12\% on one and from 36.07\% to 32.33\% on the other. In human evaluation, the resulting model also outperforms a comparable system with the same architecture trained on 98 hours of cleaner, less realistic speech. We release the data and models to support robust ASR for predominantly oral languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara
Diarra, Yacouba
Kamate, Panga Azazia
Coulibaly, Nouhoum Souleymane
Leventhal, Michael
Computation and Language
We present Kunkado, a 160-hour Bambara ASR dataset compiled from Malian radio archives to capture present-day spontaneous speech across a wide range of topics. It includes code-switching, disfluencies, background noise, and overlapping speakers that practical ASR systems encounter in real-world use. We finetuned Parakeet-based models on a 33.47-hour human-reviewed subset and apply pragmatic transcript normalization to reduce variability in number formatting, tags, and code-switching annotations. Evaluated on two real-world test sets, finetuning with Kunkado reduces WER from 44.47\% to 37.12\% on one and from 36.07\% to 32.33\% on the other. In human evaluation, the resulting model also outperforms a comparable system with the same architecture trained on 98 hours of cleaner, less realistic speech. We release the data and models to support robust ASR for predominantly oral languages.
title Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara
topic Computation and Language
url https://arxiv.org/abs/2512.19400