KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wanjawa, Barack W., Wanzare, Lilian D. A., Indede, Florence, McOnyango, Owen, Muchemi, Lawrence, Ombui, Edward
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916567578574848
author Wanjawa, Barack W.
Wanzare, Lilian D. A.
Indede, Florence
McOnyango, Owen
Muchemi, Lawrence
Ombui, Edward
author_facet Wanjawa, Barack W.
Wanzare, Lilian D. A.
Indede, Florence
McOnyango, Owen
Muchemi, Lawrence
Ombui, Edward
contents The need for Question Answering datasets in low resource languages is the motivation of this research, leading to the development of Kencorpus Swahili Question Answering Dataset, KenSwQuAD. This dataset is annotated from raw story texts of Swahili low resource language, which is a predominantly spoken in Eastern African and in other parts of the world. Question Answering (QA) datasets are important for machine comprehension of natural language for tasks such as internet search and dialog systems. Machine learning systems need training data such as the gold standard Question Answering set developed in this research. The research engaged annotators to formulate QA pairs from Swahili texts collected by the Kencorpus project, a Kenyan languages corpus. The project annotated 1,445 texts from the total 2,585 texts with at least 5 QA pairs each, resulting into a final dataset of 7,526 QA pairs. A quality assurance set of 12.5% of the annotated texts confirmed that the QA pairs were all correctly annotated. A proof of concept on applying the set to the QA task confirmed that the dataset can be usable for such tasks. KenSwQuAD has also contributed to resourcing of the Swahili language.
format Preprint
id arxiv_https___arxiv_org_abs_2205_02364
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language
Wanjawa, Barack W.
Wanzare, Lilian D. A.
Indede, Florence
McOnyango, Owen
Muchemi, Lawrence
Ombui, Edward
Computation and Language
Machine Learning
I.2.7
The need for Question Answering datasets in low resource languages is the motivation of this research, leading to the development of Kencorpus Swahili Question Answering Dataset, KenSwQuAD. This dataset is annotated from raw story texts of Swahili low resource language, which is a predominantly spoken in Eastern African and in other parts of the world. Question Answering (QA) datasets are important for machine comprehension of natural language for tasks such as internet search and dialog systems. Machine learning systems need training data such as the gold standard Question Answering set developed in this research. The research engaged annotators to formulate QA pairs from Swahili texts collected by the Kencorpus project, a Kenyan languages corpus. The project annotated 1,445 texts from the total 2,585 texts with at least 5 QA pairs each, resulting into a final dataset of 7,526 QA pairs. A quality assurance set of 12.5% of the annotated texts confirmed that the QA pairs were all correctly annotated. A proof of concept on applying the set to the QA task confirmed that the dataset can be usable for such tasks. KenSwQuAD has also contributed to resourcing of the Swahili language.
title KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language
topic Computation and Language
Machine Learning
I.2.7
url https://arxiv.org/abs/2205.02364