Correcting FLORES Evaluation Dataset for Four African Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Abdulmumin, Idris, Mkhwanazi, Sthembiso, Mbooi, Mahlatse S., Muhammad, Shamsuddeen Hassan, Ahmad, Ibrahim Said, Putini, Neo, Mathebula, Miehleketo, Shingange, Matimba, Gwadabe, Tajuddeen, Marivate, Vukosi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916424354627584
author Abdulmumin, Idris
Mkhwanazi, Sthembiso
Mbooi, Mahlatse S.
Muhammad, Shamsuddeen Hassan
Ahmad, Ibrahim Said
Putini, Neo
Mathebula, Miehleketo
Shingange, Matimba
Gwadabe, Tajuddeen
Marivate, Vukosi
author_facet Abdulmumin, Idris
Mkhwanazi, Sthembiso
Mbooi, Mahlatse S.
Muhammad, Shamsuddeen Hassan
Ahmad, Ibrahim Said
Putini, Neo
Mathebula, Miehleketo
Shingange, Matimba
Gwadabe, Tajuddeen
Marivate, Vukosi
contents This paper describes the corrections made to the FLORES evaluation (dev and devtest) dataset for four African languages, namely Hausa, Northern Sotho (Sepedi), Xitsonga, and isiZulu. The original dataset, though groundbreaking in its coverage of low-resource languages, exhibited various inconsistencies and inaccuracies in the reviewed languages that could potentially hinder the integrity of the evaluation of downstream tasks in natural language processing (NLP), especially machine translation. Through a meticulous review process by native speakers, several corrections were identified and implemented, improving the overall quality and reliability of the dataset. For each language, we provide a concise summary of the errors encountered and corrected and also present some statistical analysis that measures the difference between the existing and corrected datasets. We believe that our corrections improve the linguistic accuracy and reliability of the data and, thereby, contribute to a more effective evaluation of NLP tasks involving the four African languages. Finally, we recommend that future translation efforts, particularly in low-resource languages, prioritize the active involvement of native speakers at every stage of the process to ensure linguistic accuracy and cultural relevance.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Correcting FLORES Evaluation Dataset for Four African Languages
Abdulmumin, Idris
Mkhwanazi, Sthembiso
Mbooi, Mahlatse S.
Muhammad, Shamsuddeen Hassan
Ahmad, Ibrahim Said
Putini, Neo
Mathebula, Miehleketo
Shingange, Matimba
Gwadabe, Tajuddeen
Marivate, Vukosi
Computation and Language
This paper describes the corrections made to the FLORES evaluation (dev and devtest) dataset for four African languages, namely Hausa, Northern Sotho (Sepedi), Xitsonga, and isiZulu. The original dataset, though groundbreaking in its coverage of low-resource languages, exhibited various inconsistencies and inaccuracies in the reviewed languages that could potentially hinder the integrity of the evaluation of downstream tasks in natural language processing (NLP), especially machine translation. Through a meticulous review process by native speakers, several corrections were identified and implemented, improving the overall quality and reliability of the dataset. For each language, we provide a concise summary of the errors encountered and corrected and also present some statistical analysis that measures the difference between the existing and corrected datasets. We believe that our corrections improve the linguistic accuracy and reliability of the data and, thereby, contribute to a more effective evaluation of NLP tasks involving the four African languages. Finally, we recommend that future translation efforts, particularly in low-resource languages, prioritize the active involvement of native speakers at every stage of the process to ensure linguistic accuracy and cultural relevance.
title Correcting FLORES Evaluation Dataset for Four African Languages
topic Computation and Language
url https://arxiv.org/abs/2409.00626