Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Mingyang, Lange, Lukas, Adel, Heike, Ma, Yunpu, Strötgen, Jannik, Schütze, Hinrich
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912593799544832
author Wang, Mingyang
Lange, Lukas
Adel, Heike
Ma, Yunpu
Strötgen, Jannik
Schütze, Hinrich
author_facet Wang, Mingyang
Lange, Lukas
Adel, Heike
Ma, Yunpu
Strötgen, Jannik
Schütze, Hinrich
contents Reasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps. However, language mixing, i.e., reasoning steps containing tokens from languages other than the prompt, has been observed in their outputs and shown to affect performance, though its impact remains debated. We present the first systematic study of language mixing in RLMs, examining its patterns, impact, and internal causes across 15 languages, 7 task difficulty levels, and 18 subject areas, and show how all three factors influence language mixing. Moreover, we demonstrate that the choice of reasoning language significantly affects performance: forcing models to reason in Latin or Han scripts via constrained decoding notably improves accuracy. Finally, we show that the script composition of reasoning traces closely aligns with that of the model's internal representations, indicating that language mixing reflects latent processing preferences in RLMs. Our findings provide actionable insights for optimizing multilingual reasoning and open new directions for controlling reasoning languages to build more interpretable and adaptable RLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
Wang, Mingyang
Lange, Lukas
Adel, Heike
Ma, Yunpu
Strötgen, Jannik
Schütze, Hinrich
Computation and Language
Reasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps. However, language mixing, i.e., reasoning steps containing tokens from languages other than the prompt, has been observed in their outputs and shown to affect performance, though its impact remains debated. We present the first systematic study of language mixing in RLMs, examining its patterns, impact, and internal causes across 15 languages, 7 task difficulty levels, and 18 subject areas, and show how all three factors influence language mixing. Moreover, we demonstrate that the choice of reasoning language significantly affects performance: forcing models to reason in Latin or Han scripts via constrained decoding notably improves accuracy. Finally, we show that the script composition of reasoning traces closely aligns with that of the model's internal representations, indicating that language mixing reflects latent processing preferences in RLMs. Our findings provide actionable insights for optimizing multilingual reasoning and open new directions for controlling reasoning languages to build more interpretable and adaptable RLMs.
title Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
topic Computation and Language
url https://arxiv.org/abs/2505.14815