WolBanking77: Wolof Banking Speech Intent Classification Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kandji, Abdou Karim, Precioso, Frédéric, Ba, Cheikh, Ndiaye, Samba, Ndione, Augustin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911230770282496
author Kandji, Abdou Karim
Precioso, Frédéric
Ba, Cheikh
Ndiaye, Samba
Ndione, Augustin
author_facet Kandji, Abdou Karim
Precioso, Frédéric
Ba, Cheikh
Ndiaye, Samba
Ndione, Augustin
contents Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90\% of the population, while the national illiteracy rate remains at of 42\%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification. WolBanking77 currently contains 9,791 text sentences in the banking domain and more than 4 hours of spoken sentences. Experiments on various baselines are conducted in this work, including text and voice state-of-the-art models. The results are very promising on this current dataset. In addition, this paper presents an in-depth examination of the dataset's contents. We report baseline F1-scores and word error rates metrics respectively on NLP and ASR models trained on WolBanking77 dataset and also comparisons between models. Dataset and code available at: https://github.com/abdoukarim/wolbanking77.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19271
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WolBanking77: Wolof Banking Speech Intent Classification Dataset
Kandji, Abdou Karim
Precioso, Frédéric
Ba, Cheikh
Ndiaye, Samba
Ndione, Augustin
Computation and Language
Artificial Intelligence
Machine Learning
Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90\% of the population, while the national illiteracy rate remains at of 42\%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification. WolBanking77 currently contains 9,791 text sentences in the banking domain and more than 4 hours of spoken sentences. Experiments on various baselines are conducted in this work, including text and voice state-of-the-art models. The results are very promising on this current dataset. In addition, this paper presents an in-depth examination of the dataset's contents. We report baseline F1-scores and word error rates metrics respectively on NLP and ASR models trained on WolBanking77 dataset and also comparisons between models. Dataset and code available at: https://github.com/abdoukarim/wolbanking77.
title WolBanking77: Wolof Banking Speech Intent Classification Dataset
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.19271