Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zou, Wei, Yang, Sen, Bao, Yu, Huang, Shujian, Chen, Jiajun, Cheng, Shanbo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908368249028608
author Zou, Wei
Yang, Sen
Bao, Yu
Huang, Shujian
Chen, Jiajun
Cheng, Shanbo
author_facet Zou, Wei
Yang, Sen
Bao, Yu
Huang, Shujian
Chen, Jiajun
Cheng, Shanbo
contents The rise of Large Language Models (LLMs) has reshaped machine translation (MT), but multilingual MT still relies heavily on parallel data for supervised fine-tuning (SFT), facing challenges like data scarcity for low-resource languages and catastrophic forgetting. To address these issues, we propose TRANS-ZERO, a self-play framework that leverages only monolingual data and the intrinsic multilingual knowledge of LLM. TRANS-ZERO combines Genetic Monte-Carlo Tree Search (G-MCTS) with preference optimization, achieving strong translation performance that rivals supervised methods. Experiments demonstrate that this approach not only matches the performance of models trained on large-scale parallel data but also excels in non-English translation directions. Further analysis reveals that G-MCTS itself significantly enhances translation quality by exploring semantically consistent candidates through iterative translations, providing a robust foundation for the framework's succuss.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14669
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
Zou, Wei
Yang, Sen
Bao, Yu
Huang, Shujian
Chen, Jiajun
Cheng, Shanbo
Computation and Language
The rise of Large Language Models (LLMs) has reshaped machine translation (MT), but multilingual MT still relies heavily on parallel data for supervised fine-tuning (SFT), facing challenges like data scarcity for low-resource languages and catastrophic forgetting. To address these issues, we propose TRANS-ZERO, a self-play framework that leverages only monolingual data and the intrinsic multilingual knowledge of LLM. TRANS-ZERO combines Genetic Monte-Carlo Tree Search (G-MCTS) with preference optimization, achieving strong translation performance that rivals supervised methods. Experiments demonstrate that this approach not only matches the performance of models trained on large-scale parallel data but also excels in non-English translation directions. Further analysis reveals that G-MCTS itself significantly enhances translation quality by exploring semantically consistent candidates through iterative translations, providing a robust foundation for the framework's succuss.
title Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
topic Computation and Language
url https://arxiv.org/abs/2504.14669