EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tonja, Atnafu Lambebo, Kolesnikova, Olga, Gelbukh, Alexander, Kalita, Jugal
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917624801132544
author Tonja, Atnafu Lambebo
Kolesnikova, Olga
Gelbukh, Alexander
Kalita, Jugal
author_facet Tonja, Atnafu Lambebo
Kolesnikova, Olga
Gelbukh, Alexander
Kalita, Jugal
contents Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for low-resource languages. This is due to the smaller size of available parallel corpora in these languages, if such corpora are available at all. NLP in Ethiopian languages suffers from the same issues due to the unavailability of publicly accessible datasets for NLP tasks, including MT. To help the research community and foster research for Ethiopian languages, we introduce EthioMT -- a new parallel corpus for 15 languages. We also create a new benchmark by collecting a dataset for better-researched languages in Ethiopia. We evaluate the newly collected corpus and the benchmark dataset for 23 Ethiopian languages using transformer and fine-tuning approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EthioMT: Parallel Corpus for Low-resource Ethiopian Languages
Tonja, Atnafu Lambebo
Kolesnikova, Olga
Gelbukh, Alexander
Kalita, Jugal
Computation and Language
Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for low-resource languages. This is due to the smaller size of available parallel corpora in these languages, if such corpora are available at all. NLP in Ethiopian languages suffers from the same issues due to the unavailability of publicly accessible datasets for NLP tasks, including MT. To help the research community and foster research for Ethiopian languages, we introduce EthioMT -- a new parallel corpus for 15 languages. We also create a new benchmark by collecting a dataset for better-researched languages in Ethiopia. We evaluate the newly collected corpus and the benchmark dataset for 23 Ethiopian languages using transformer and fine-tuning approaches.
title EthioMT: Parallel Corpus for Low-resource Ethiopian Languages
topic Computation and Language
url https://arxiv.org/abs/2403.19365