Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kalejaiye, Oluwadara, Beyene, Luel Hagos, Adelani, David Ifeoluwa, Edet, Mmekut-Mfon Gabriel, Akpan, Aniefon Daniel, Urua, Eno-Abasi, Andy, Anietie
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915608550965248
author Kalejaiye, Oluwadara
Beyene, Luel Hagos
Adelani, David Ifeoluwa
Edet, Mmekut-Mfon Gabriel
Akpan, Aniefon Daniel
Urua, Eno-Abasi
Andy, Anietie
author_facet Kalejaiye, Oluwadara
Beyene, Luel Hagos
Adelani, David Ifeoluwa
Edet, Mmekut-Mfon Gabriel
Akpan, Aniefon Daniel
Urua, Eno-Abasi
Andy, Anietie
contents Nigeria is the most populous country in Africa with a population of more than 200 million people. More than 500 languages are spoken in Nigeria and it is one of the most linguistically diverse countries in the world. Despite this, natural language processing (NLP) research has mostly focused on the following four languages: Hausa, Igbo, Nigerian-Pidgin, and Yoruba (i.e <1% of the languages spoken in Nigeria). This is in part due to the unavailability of textual data in these languages to train and apply NLP algorithms. In this work, we introduce ibom -- a dataset for machine translation and topic classification in four Coastal Nigerian languages from the Akwa Ibom State region: Anaang, Efik, Ibibio, and Oro. These languages are not represented in Google Translate or in major benchmarks such as Flores-200 or SIB-200. We focus on extending Flores-200 benchmark to these languages, and further align the translated texts with topic labels based on SIB-200 classification dataset. Our evaluation shows that current LLMs perform poorly on machine translation for these languages in both zero-and-few shot settings. However, we find the few-shot samples to steadily improve topic classification with more shots.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06531
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages
Kalejaiye, Oluwadara
Beyene, Luel Hagos
Adelani, David Ifeoluwa
Edet, Mmekut-Mfon Gabriel
Akpan, Aniefon Daniel
Urua, Eno-Abasi
Andy, Anietie
Computation and Language
Artificial Intelligence
Nigeria is the most populous country in Africa with a population of more than 200 million people. More than 500 languages are spoken in Nigeria and it is one of the most linguistically diverse countries in the world. Despite this, natural language processing (NLP) research has mostly focused on the following four languages: Hausa, Igbo, Nigerian-Pidgin, and Yoruba (i.e <1% of the languages spoken in Nigeria). This is in part due to the unavailability of textual data in these languages to train and apply NLP algorithms. In this work, we introduce ibom -- a dataset for machine translation and topic classification in four Coastal Nigerian languages from the Akwa Ibom State region: Anaang, Efik, Ibibio, and Oro. These languages are not represented in Google Translate or in major benchmarks such as Flores-200 or SIB-200. We focus on extending Flores-200 benchmark to these languages, and further align the translated texts with topic labels based on SIB-200 classification dataset. Our evaluation shows that current LLMs perform poorly on machine translation for these languages in both zero-and-few shot settings. However, we find the few-shot samples to steadily improve topic classification with more shots.
title Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.06531