From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Yingli, Lai, Wen, Wang, Shuo, Gao, Ge, Luo, Kangyang, Fraser, Alexander, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911222321905664
author Shen, Yingli
Lai, Wen
Wang, Shuo
Gao, Ge
Luo, Kangyang
Fraser, Alexander
Sun, Maosong
author_facet Shen, Yingli
Lai, Wen
Wang, Shuo
Gao, Ge
Luo, Kangyang
Fraser, Alexander
Sun, Maosong
contents Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics. In contrast, multi-way parallel data, where identical content is aligned across multiple languages, provides stronger cross-lingual consistency and offers greater potential for improving multilingual performance. In this paper, we introduce a large-scale, high-quality multi-way parallel corpus, TED2025, based on TED Talks. The corpus spans 113 languages, with up to 50 languages aligned in parallel, ensuring extensive multilingual coverage. Using this dataset, we investigate best practices for leveraging multi-way parallel data to enhance LLMs, including strategies for continued pretraining, instruction tuning, and the analysis of key influencing factors. Experiments on six multilingual benchmarks show that models trained on multiway parallel data consistently outperform those trained on unaligned multilingual data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
Shen, Yingli
Lai, Wen
Wang, Shuo
Gao, Ge
Luo, Kangyang
Fraser, Alexander
Sun, Maosong
Computation and Language
Artificial Intelligence
Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics. In contrast, multi-way parallel data, where identical content is aligned across multiple languages, provides stronger cross-lingual consistency and offers greater potential for improving multilingual performance. In this paper, we introduce a large-scale, high-quality multi-way parallel corpus, TED2025, based on TED Talks. The corpus spans 113 languages, with up to 50 languages aligned in parallel, ensuring extensive multilingual coverage. Using this dataset, we investigate best practices for leveraging multi-way parallel data to enhance LLMs, including strategies for continued pretraining, instruction tuning, and the analysis of key influencing factors. Experiments on six multilingual benchmarks show that models trained on multiway parallel data consistently outperform those trained on unaligned multilingual data.
title From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.14045