MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kot, Artur, Koszowski, Mikołaj, Chojnowski, Wojciech, Rutkowski, Mieszko, Nowakowski, Artur, Guttmann, Kamil, Pokrywka, Mikołaj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929722325204992
author Kot, Artur
Koszowski, Mikołaj
Chojnowski, Wojciech
Rutkowski, Mieszko
Nowakowski, Artur
Guttmann, Kamil
Pokrywka, Mikołaj
author_facet Kot, Artur
Koszowski, Mikołaj
Chojnowski, Wojciech
Rutkowski, Mieszko
Nowakowski, Artur
Guttmann, Kamil
Pokrywka, Mikołaj
contents Does multilingual Neural Machine Translation (NMT) lead to The Curse of the Multlinguality or provides the Cross-lingual Knowledge Transfer within a language family? In this study, we explore multiple approaches for extending the available data-regime in NMT and we prove cross-lingual benefits even in 0-shot translation regime for low-resource languages. With this paper, we provide state-of-the-art open-source NMT models for translating between selected Slavic languages. We released our models on the HuggingFace Hub (https://hf.co/collections/allegro/multislav-6793d6b6419e5963e759a683) under the CC BY 4.0 license. Slavic language family comprises morphologically rich Central and Eastern European languages. Although counting hundreds of millions of native speakers, Slavic Neural Machine Translation is under-studied in our opinion. Recently, most NMT research focuses either on: high-resource languages like English, Spanish, and German - in WMT23 General Translation Task 7 out of 8 task directions are from or to English; massively multilingual models covering multiple language groups; or evaluation techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality
Kot, Artur
Koszowski, Mikołaj
Chojnowski, Wojciech
Rutkowski, Mieszko
Nowakowski, Artur
Guttmann, Kamil
Pokrywka, Mikołaj
Computation and Language
Does multilingual Neural Machine Translation (NMT) lead to The Curse of the Multlinguality or provides the Cross-lingual Knowledge Transfer within a language family? In this study, we explore multiple approaches for extending the available data-regime in NMT and we prove cross-lingual benefits even in 0-shot translation regime for low-resource languages. With this paper, we provide state-of-the-art open-source NMT models for translating between selected Slavic languages. We released our models on the HuggingFace Hub (https://hf.co/collections/allegro/multislav-6793d6b6419e5963e759a683) under the CC BY 4.0 license. Slavic language family comprises morphologically rich Central and Eastern European languages. Although counting hundreds of millions of native speakers, Slavic Neural Machine Translation is under-studied in our opinion. Recently, most NMT research focuses either on: high-resource languages like English, Spanish, and German - in WMT23 General Translation Task 7 out of 8 task directions are from or to English; massively multilingual models covering multiple language groups; or evaluation techniques.
title MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality
topic Computation and Language
url https://arxiv.org/abs/2502.14509