Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Elhady, Ahmed, Agirre, Eneko, Artetxe, Mikel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910279000915968
author Elhady, Ahmed
Agirre, Eneko
Artetxe, Mikel
author_facet Elhady, Ahmed
Agirre, Eneko
Artetxe, Mikel
contents Despite expanding their multilingual coverage, the advanced reasoning capabilities of LLMs remain largely confined to a few high-resource languages like English. To address this, we propose an unsupervised Reinforcement Learning (RL) approach to enhance multilingual reasoning by enforcing cross-lingual self-consistency: the principle that a model should produce the same final answer for equivalent problems in different languages. Existing methods are limited by the scarcity of multilingual reasoning data and show weak generalization to unseen languages. Our approach requires neither gold answers nor parallel data, and it achieves average gains of up to 21.7% on MGSM across 10 languages. In addition, our method demonstrates strong generalization, with an 18.2% mean improvement on MGSM languages unseen during training, and up to 6.2% gain on 3 out-of-distribution benchmarks. These results show the potential of consistency-based methods to improve the multilingual capabilities of LLMs without requiring supervised data.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01464
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models
Elhady, Ahmed
Agirre, Eneko
Artetxe, Mikel
Computation and Language
Despite expanding their multilingual coverage, the advanced reasoning capabilities of LLMs remain largely confined to a few high-resource languages like English. To address this, we propose an unsupervised Reinforcement Learning (RL) approach to enhance multilingual reasoning by enforcing cross-lingual self-consistency: the principle that a model should produce the same final answer for equivalent problems in different languages. Existing methods are limited by the scarcity of multilingual reasoning data and show weak generalization to unseen languages. Our approach requires neither gold answers nor parallel data, and it achieves average gains of up to 21.7% on MGSM across 10 languages. In addition, our method demonstrates strong generalization, with an 18.2% mean improvement on MGSM languages unseen during training, and up to 6.2% gain on 3 out-of-distribution benchmarks. These results show the potential of consistency-based methods to improve the multilingual capabilities of LLMs without requiring supervised data.
title Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models
topic Computation and Language
url https://arxiv.org/abs/2606.01464