TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Xu, Lai, Zhejian, Huang, Zixian, Chen, Jiajun, Huang, Shujian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908915664420864
author Huang, Xu
Lai, Zhejian
Huang, Zixian
Chen, Jiajun
Huang, Shujian
author_facet Huang, Xu
Lai, Zhejian
Huang, Zixian
Chen, Jiajun
Huang, Shujian
contents Large Language Models (LLMs) have demonstrated remarkable proficiency in English mathematical reasoning, yet a significant performance disparity persists in multilingual contexts, largely attributed to deficiencies in language understanding. To bridge this gap, we introduce Translation-Augmented Policy Optimization (TAPO), a novel reinforcement learning framework built upon GRPO. TAPO enforces an explicit alignment strategy where the model leverages English as a pivot and follows an understand-then-reason paradigm. Crucially, we employ a step-level relative advantage mechanism that decouples understanding from reasoning, allowing the integration of translation quality rewards without introducing optimization conflicts. Extensive experiments reveal that TAPO effectively synergizes language understanding with reasoning capabilities and is compatible with various models. It outperforms baseline methods in both multilingual mathematical reasoning and translation tasks, while generalizing well to unseen languages and out-of-domain tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25419
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning
Huang, Xu
Lai, Zhejian
Huang, Zixian
Chen, Jiajun
Huang, Shujian
Computation and Language
Large Language Models (LLMs) have demonstrated remarkable proficiency in English mathematical reasoning, yet a significant performance disparity persists in multilingual contexts, largely attributed to deficiencies in language understanding. To bridge this gap, we introduce Translation-Augmented Policy Optimization (TAPO), a novel reinforcement learning framework built upon GRPO. TAPO enforces an explicit alignment strategy where the model leverages English as a pivot and follows an understand-then-reason paradigm. Crucially, we employ a step-level relative advantage mechanism that decouples understanding from reasoning, allowing the integration of translation quality rewards without introducing optimization conflicts. Extensive experiments reveal that TAPO effectively synergizes language understanding with reasoning capabilities and is compatible with various models. It outperforms baseline methods in both multilingual mathematical reasoning and translation tasks, while generalizing well to unseen languages and out-of-domain tasks.
title TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning
topic Computation and Language
url https://arxiv.org/abs/2603.25419