Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Menglong, Gao, Pengzhi, Liu, Wei, Luan, Jian, Wang, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929728478248960
author Cui, Menglong
Gao, Pengzhi
Liu, Wei
Luan, Jian
Wang, Bin
author_facet Cui, Menglong
Gao, Pengzhi
Liu, Wei
Luan, Jian
Wang, Bin
contents Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. In this paper, we systematically explore the abilities of open LLMs with less than ten billion parameters to handle multilingual machine translation (MT) tasks. We conduct comprehensive evaluations on six popular LLMs and find that models like Gemma2-9B exhibit impressive multilingual translation capabilities. We then introduce the Parallel-First Monolingual-Second (PFMS) data mixing strategy in the continual pretraining stage to further enhance the MT performance and present GemmaX2-28, a 9B model achieving top-tier multilingual translation performance across 28 languages. Specifically, GemmaX2-28 consistently outperforms the state-of-the-art (SOTA) models such as TowerInstruct and XALMA and achieves competitive performance with Google Translate and GPT-4-turbo.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02481
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
Cui, Menglong
Gao, Pengzhi
Liu, Wei
Luan, Jian
Wang, Bin
Computation and Language
Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. In this paper, we systematically explore the abilities of open LLMs with less than ten billion parameters to handle multilingual machine translation (MT) tasks. We conduct comprehensive evaluations on six popular LLMs and find that models like Gemma2-9B exhibit impressive multilingual translation capabilities. We then introduce the Parallel-First Monolingual-Second (PFMS) data mixing strategy in the continual pretraining stage to further enhance the MT performance and present GemmaX2-28, a 9B model achieving top-tier multilingual translation performance across 28 languages. Specifically, GemmaX2-28 consistently outperforms the state-of-the-art (SOTA) models such as TowerInstruct and XALMA and achieves competitive performance with Google Translate and GPT-4-turbo.
title Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
topic Computation and Language
url https://arxiv.org/abs/2502.02481