MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: She, Shuaijie, Zou, Wei, Huang, Shujian, Zhu, Wenhao, Liu, Xiang, Geng, Xiang, Chen, Jiajun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909168839950336
author She, Shuaijie
Zou, Wei
Huang, Shujian
Zhu, Wenhao
Liu, Xiang
Geng, Xiang
Chen, Jiajun
author_facet She, Shuaijie
Zou, Wei
Huang, Shujian
Zhu, Wenhao
Liu, Xiang
Geng, Xiang
Chen, Jiajun
contents Though reasoning abilities are considered language-agnostic, existing LLMs exhibit inconsistent reasoning abilities across different languages, e.g., reasoning in the dominant language like English is superior to other languages due to the imbalance of multilingual training data. To enhance reasoning abilities in non-dominant languages, we propose a Multilingual-Alignment-as-Preference Optimization framework (MAPO), aiming to align the reasoning processes in other languages with the dominant language. Specifically, we harness an off-the-shelf translation model for the consistency between answers in non-dominant and dominant languages, which we adopt as the preference for optimization, e.g., Direct Preference Optimization (DPO) or Proximal Policy Optimization (PPO). Experiments show that MAPO stably achieves significant improvements in the multilingual reasoning of various models on all three benchmarks (MSVAMP +16.2%, MGSM +6.1%, and MNumGLUESub +13.3%), with improved reasoning consistency across languages.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization
She, Shuaijie
Zou, Wei
Huang, Shujian
Zhu, Wenhao
Liu, Xiang
Geng, Xiang
Chen, Jiajun
Computation and Language
Though reasoning abilities are considered language-agnostic, existing LLMs exhibit inconsistent reasoning abilities across different languages, e.g., reasoning in the dominant language like English is superior to other languages due to the imbalance of multilingual training data. To enhance reasoning abilities in non-dominant languages, we propose a Multilingual-Alignment-as-Preference Optimization framework (MAPO), aiming to align the reasoning processes in other languages with the dominant language. Specifically, we harness an off-the-shelf translation model for the consistency between answers in non-dominant and dominant languages, which we adopt as the preference for optimization, e.g., Direct Preference Optimization (DPO) or Proximal Policy Optimization (PPO). Experiments show that MAPO stably achieves significant improvements in the multilingual reasoning of various models on all three benchmarks (MSVAMP +16.2%, MGSM +6.1%, and MNumGLUESub +13.3%), with improved reasoning consistency across languages.
title MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization
topic Computation and Language
url https://arxiv.org/abs/2401.06838