Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911114443358208 |
|---|---|
| author | Cheng, Shanbo Bao, Yu Cao, Qian Huang, Luyang Kang, Liyan Liu, Zhicheng Lu, Yu Zhu, Wenhao Chen, Jingwen Huang, Zhichao Li, Tao Li, Yifu Lin, Huiying Liu, Sitong Peng, Ningxin She, Shuaijie Xu, Lu Xu, Nuo Yang, Sen Yu, Runsheng Yu, Yiming Zou, Liehao Li, Hang Lu, Lu Wang, Yuxuan Wu, Yonghui |
| author_facet | Cheng, Shanbo Bao, Yu Cao, Qian Huang, Luyang Kang, Liyan Liu, Zhicheng Lu, Yu Zhu, Wenhao Chen, Jingwen Huang, Zhichao Li, Tao Li, Yifu Lin, Huiying Liu, Sitong Peng, Ningxin She, Shuaijie Xu, Lu Xu, Nuo Yang, Sen Yu, Runsheng Yu, Yiming Zou, Liehao Li, Hang Lu, Lu Wang, Yuxuan Wu, Yonghui |
| contents | Multilingual translation stands as a challenging task for large language models (LLMs) to handle intricate language patterns and stilted translations that arise in automated translations. In this paper, we introduce Seed-X, a family of open-source LLMs comprising instruct and reasoning models, pushing the limits of translation capability with 7B parameter size. The base model is pre-trained on a diverse, high-quality dataset encompassing both monolingual and bilingual content across 28 languages, harnessing the full potential of multilingual data. The instruct model is then finetuned to translate by Chain-of-Thought (CoT) reasoning and further enhanced through reinforcement learning (RL) to achieve better generalization across diverse language pairs. Seed-X achieves performance comparable to leading closed-source models, including Gemini-2.5 and GPT-4o, across 28 languages, and significantly outperforms larger open-source models in both automatic metrics and human evaluations. We share the best practices through our optimization process, and make the parameter public available for advancing translation research and applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_13618 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters Cheng, Shanbo Bao, Yu Cao, Qian Huang, Luyang Kang, Liyan Liu, Zhicheng Lu, Yu Zhu, Wenhao Chen, Jingwen Huang, Zhichao Li, Tao Li, Yifu Lin, Huiying Liu, Sitong Peng, Ningxin She, Shuaijie Xu, Lu Xu, Nuo Yang, Sen Yu, Runsheng Yu, Yiming Zou, Liehao Li, Hang Lu, Lu Wang, Yuxuan Wu, Yonghui Computation and Language Artificial Intelligence Multilingual translation stands as a challenging task for large language models (LLMs) to handle intricate language patterns and stilted translations that arise in automated translations. In this paper, we introduce Seed-X, a family of open-source LLMs comprising instruct and reasoning models, pushing the limits of translation capability with 7B parameter size. The base model is pre-trained on a diverse, high-quality dataset encompassing both monolingual and bilingual content across 28 languages, harnessing the full potential of multilingual data. The instruct model is then finetuned to translate by Chain-of-Thought (CoT) reasoning and further enhanced through reinforcement learning (RL) to achieve better generalization across diverse language pairs. Seed-X achieves performance comparable to leading closed-source models, including Gemini-2.5 and GPT-4o, across 28 languages, and significantly outperforms larger open-source models in both automatic metrics and human evaluations. We share the best practices through our optimization process, and make the parameter public available for advancing translation research and applications. |
| title | Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2507.13618 |