Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ahia, Orevaoghene, Aremu, Anuoluwapo, Abagyan, Diana, Gonen, Hila, Adelani, David Ifeoluwa, Abolade, Daud, Smith, Noah A., Tsvetkov, Yulia
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913407042584576
author Ahia, Orevaoghene
Aremu, Anuoluwapo
Abagyan, Diana
Gonen, Hila
Adelani, David Ifeoluwa
Abolade, Daud
Smith, Noah A.
Tsvetkov, Yulia
author_facet Ahia, Orevaoghene
Aremu, Anuoluwapo
Abagyan, Diana
Gonen, Hila
Adelani, David Ifeoluwa
Abolade, Daud
Smith, Noah A.
Tsvetkov, Yulia
contents Yorùbá an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in disparities for dialects and varieties for which there are little to no resources or tools. We take steps towards bridging this gap by introducing a new high-quality parallel text and speech corpus YORÙLECT across three domains and four regional Yorùbá dialects. To develop this corpus, we engaged native speakers, travelling to communities where these dialects are spoken, to collect text and speech data. Using our newly created corpus, we conducted extensive experiments on (text) machine translation, automatic speech recognition, and speech-to-text translation. Our results reveal substantial performance disparities between standard Yorùbá and the other dialects across all tasks. However, we also show that with dialect-adaptive finetuning, we are able to narrow this gap. We believe our dataset and experimental analysis will contribute greatly to developing NLP tools for Yorùbá and its dialects, and potentially for other African languages, by improving our understanding of existing challenges and offering a high-quality dataset for further development. We release YORÙLECT dataset and models publicly under an open license.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19564
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects
Ahia, Orevaoghene
Aremu, Anuoluwapo
Abagyan, Diana
Gonen, Hila
Adelani, David Ifeoluwa
Abolade, Daud
Smith, Noah A.
Tsvetkov, Yulia
Computation and Language
Yorùbá an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in disparities for dialects and varieties for which there are little to no resources or tools. We take steps towards bridging this gap by introducing a new high-quality parallel text and speech corpus YORÙLECT across three domains and four regional Yorùbá dialects. To develop this corpus, we engaged native speakers, travelling to communities where these dialects are spoken, to collect text and speech data. Using our newly created corpus, we conducted extensive experiments on (text) machine translation, automatic speech recognition, and speech-to-text translation. Our results reveal substantial performance disparities between standard Yorùbá and the other dialects across all tasks. However, we also show that with dialect-adaptive finetuning, we are able to narrow this gap. We believe our dataset and experimental analysis will contribute greatly to developing NLP tools for Yorùbá and its dialects, and potentially for other African languages, by improving our understanding of existing challenges and offering a high-quality dataset for further development. We release YORÙLECT dataset and models publicly under an open license.
title Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects
topic Computation and Language
url https://arxiv.org/abs/2406.19564