R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Minggui, Liu, Yilun, Tao, Shimin, Luo, Yuanchang, Zeng, Hongyong, Su, Chang, Zhang, Li, Ma, Hongxia, Wei, Daimeng, Meng, Weibin, Yang, Hao, Chen, Boxing, Yoshie, Osamu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910968363089920
author He, Minggui
Liu, Yilun
Tao, Shimin
Luo, Yuanchang
Zeng, Hongyong
Su, Chang
Zhang, Li
Ma, Hongxia
Wei, Daimeng
Meng, Weibin
Yang, Hao
Chen, Boxing
Yoshie, Osamu
author_facet He, Minggui
Liu, Yilun
Tao, Shimin
Luo, Yuanchang
Zeng, Hongyong
Su, Chang
Zhang, Li
Ma, Hongxia
Wei, Daimeng
Meng, Weibin
Yang, Hao
Chen, Boxing
Yoshie, Osamu
contents Despite recent breakthroughs in reasoning-enhanced large language models (LLMs) like DeepSeek-R1, incorporating inference-time reasoning into machine translation (MT), where human translators naturally employ structured, multi-layered reasoning chain-of-thoughts (CoTs), is yet underexplored. Existing methods either design a fixed CoT tailored for a specific MT sub-task (e.g., literature translation), or rely on synthesizing CoTs unaligned with humans and supervised fine-tuning (SFT) prone to overfitting, limiting their adaptability to diverse translation scenarios. This paper introduces R1-Translator (R1-T1), a novel framework to achieve inference-time reasoning for general MT via reinforcement learning (RL) with human-aligned CoTs comprising six common patterns. Our approach pioneers three innovations: (1) extending reasoning-based translation to broader MT scenarios (e.g., multilingual MT, domain MT) unseen in the training phase; (2) formalizing six expert-curated CoT templates that mirror hybrid human strategies like context-aware paraphrasing and back translation; and (3) enabling self-evolving CoT discovery through RL. Both human and automatic evaluation results indicate a steady translation performance improvement in a total of 10+ languages and 40+ translation directions on Flores-101 test set and four domain-specific MT tasks, especially on the languages unseen from training.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning
He, Minggui
Liu, Yilun
Tao, Shimin
Luo, Yuanchang
Zeng, Hongyong
Su, Chang
Zhang, Li
Ma, Hongxia
Wei, Daimeng
Meng, Weibin
Yang, Hao
Chen, Boxing
Yoshie, Osamu
Computation and Language
Despite recent breakthroughs in reasoning-enhanced large language models (LLMs) like DeepSeek-R1, incorporating inference-time reasoning into machine translation (MT), where human translators naturally employ structured, multi-layered reasoning chain-of-thoughts (CoTs), is yet underexplored. Existing methods either design a fixed CoT tailored for a specific MT sub-task (e.g., literature translation), or rely on synthesizing CoTs unaligned with humans and supervised fine-tuning (SFT) prone to overfitting, limiting their adaptability to diverse translation scenarios. This paper introduces R1-Translator (R1-T1), a novel framework to achieve inference-time reasoning for general MT via reinforcement learning (RL) with human-aligned CoTs comprising six common patterns. Our approach pioneers three innovations: (1) extending reasoning-based translation to broader MT scenarios (e.g., multilingual MT, domain MT) unseen in the training phase; (2) formalizing six expert-curated CoT templates that mirror hybrid human strategies like context-aware paraphrasing and back translation; and (3) enabling self-evolving CoT discovery through RL. Both human and automatic evaluation results indicate a steady translation performance improvement in a total of 10+ languages and 40+ translation directions on Flores-101 test set and four domain-specific MT tasks, especially on the languages unseen from training.
title R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning
topic Computation and Language
url https://arxiv.org/abs/2502.19735