Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Burda-Lassen, Olena
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929540812505088
author Burda-Lassen, Olena
author_facet Burda-Lassen, Olena
contents Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated texts is very sparse, which limits access to knowledge about cultural traditions and customs. We have created a new Ukrainian-To-English parallel corpus of familiar Ukrainian folktales based on available English translations and suggested several new ones. We offer a combined domain-specific approach to building and augmenting this corpus, considering the nature of the domain and differences in the purpose of human versus machine translation. Our corpus is word and sentence-aligned, allowing for the best curation of meaning, specifically tailored for use as training data for machine translation models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10063
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
Burda-Lassen, Olena
Computation and Language
Artificial Intelligence
Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated texts is very sparse, which limits access to knowledge about cultural traditions and customs. We have created a new Ukrainian-To-English parallel corpus of familiar Ukrainian folktales based on available English translations and suggested several new ones. We offer a combined domain-specific approach to building and augmenting this corpus, considering the nature of the domain and differences in the purpose of human versus machine translation. Our corpus is word and sentence-aligned, allowing for the best curation of meaning, specifically tailored for use as training data for machine translation models.
title Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.10063