MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghatage, Ruturaj, Kulkarni, Aditya, Patil, Rajlaxmi, Endait, Sharvi, Joshi, Raviraj
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917645762166784
author Ghatage, Ruturaj
Kulkarni, Aditya
Patil, Rajlaxmi
Endait, Sharvi
Joshi, Raviraj
author_facet Ghatage, Ruturaj
Kulkarni, Aditya
Patil, Rajlaxmi
Endait, Sharvi
Joshi, Raviraj
contents Question-answering systems have revolutionized information retrieval, but linguistic and cultural boundaries limit their widespread accessibility. This research endeavors to bridge the gap of the absence of efficient QnA datasets in low-resource languages by translating the English Question Answering Dataset (SQuAD) using a robust data curation approach. We introduce MahaSQuAD, the first-ever full SQuAD dataset for the Indic language Marathi, consisting of 118,516 training, 11,873 validation, and 11,803 test samples. We also present a gold test set of manually verified 500 examples. Challenges in maintaining context and handling linguistic nuances are addressed, ensuring accurate translations. Moreover, as a QnA dataset cannot be simply converted into any low-resource language using translation, we need a robust method to map the answer translation to its span in the translated passage. Hence, to address this challenge, we also present a generic approach for translating SQuAD into any low-resource language. Thus, we offer a scalable approach to bridge linguistic and cultural gaps present in low-resource languages, in the realm of question-answering systems. The datasets and models are shared publicly at https://github.com/l3cube-pune/MarathiNLP .
format Preprint
id arxiv_https___arxiv_org_abs_2404_13364
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering
Ghatage, Ruturaj
Kulkarni, Aditya
Patil, Rajlaxmi
Endait, Sharvi
Joshi, Raviraj
Computation and Language
Machine Learning
Question-answering systems have revolutionized information retrieval, but linguistic and cultural boundaries limit their widespread accessibility. This research endeavors to bridge the gap of the absence of efficient QnA datasets in low-resource languages by translating the English Question Answering Dataset (SQuAD) using a robust data curation approach. We introduce MahaSQuAD, the first-ever full SQuAD dataset for the Indic language Marathi, consisting of 118,516 training, 11,873 validation, and 11,803 test samples. We also present a gold test set of manually verified 500 examples. Challenges in maintaining context and handling linguistic nuances are addressed, ensuring accurate translations. Moreover, as a QnA dataset cannot be simply converted into any low-resource language using translation, we need a robust method to map the answer translation to its span in the translated passage. Hence, to address this challenge, we also present a generic approach for translating SQuAD into any low-resource language. Thus, we offer a scalable approach to bridge linguistic and cultural gaps present in low-resource languages, in the realm of question-answering systems. The datasets and models are shared publicly at https://github.com/l3cube-pune/MarathiNLP .
title MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2404.13364