MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929722325204992 |
|---|---|
| author | Kot, Artur Koszowski, Mikołaj Chojnowski, Wojciech Rutkowski, Mieszko Nowakowski, Artur Guttmann, Kamil Pokrywka, Mikołaj |
| author_facet | Kot, Artur Koszowski, Mikołaj Chojnowski, Wojciech Rutkowski, Mieszko Nowakowski, Artur Guttmann, Kamil Pokrywka, Mikołaj |
| contents | Does multilingual Neural Machine Translation (NMT) lead to The Curse of the Multlinguality or provides the Cross-lingual Knowledge Transfer within a language family? In this study, we explore multiple approaches for extending the available data-regime in NMT and we prove cross-lingual benefits even in 0-shot translation regime for low-resource languages. With this paper, we provide state-of-the-art open-source NMT models for translating between selected Slavic languages. We released our models on the HuggingFace Hub (https://hf.co/collections/allegro/multislav-6793d6b6419e5963e759a683) under the CC BY 4.0 license. Slavic language family comprises morphologically rich Central and Eastern European languages. Although counting hundreds of millions of native speakers, Slavic Neural Machine Translation is under-studied in our opinion. Recently, most NMT research focuses either on: high-resource languages like English, Spanish, and German - in WMT23 General Translation Task 7 out of 8 task directions are from or to English; massively multilingual models covering multiple language groups; or evaluation techniques. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_14509 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality Kot, Artur Koszowski, Mikołaj Chojnowski, Wojciech Rutkowski, Mieszko Nowakowski, Artur Guttmann, Kamil Pokrywka, Mikołaj Computation and Language Does multilingual Neural Machine Translation (NMT) lead to The Curse of the Multlinguality or provides the Cross-lingual Knowledge Transfer within a language family? In this study, we explore multiple approaches for extending the available data-regime in NMT and we prove cross-lingual benefits even in 0-shot translation regime for low-resource languages. With this paper, we provide state-of-the-art open-source NMT models for translating between selected Slavic languages. We released our models on the HuggingFace Hub (https://hf.co/collections/allegro/multislav-6793d6b6419e5963e759a683) under the CC BY 4.0 license. Slavic language family comprises morphologically rich Central and Eastern European languages. Although counting hundreds of millions of native speakers, Slavic Neural Machine Translation is under-studied in our opinion. Recently, most NMT research focuses either on: high-resource languages like English, Spanish, and German - in WMT23 General Translation Task 7 out of 8 task directions are from or to English; massively multilingual models covering multiple language groups; or evaluation techniques. |
| title | MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2502.14509 |