Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910964210728960 |
|---|---|
| author | Tam, Zhi Rui Wu, Cheng-Kuang Chiu, Yu Ying Lin, Chieh-Yen Chen, Yun-Nung Lee, Hung-yi |
| author_facet | Tam, Zhi Rui Wu, Cheng-Kuang Chiu, Yu Ying Lin, Chieh-Yen Chen, Yun-Nung Lee, Hung-yi |
| contents | Large reasoning models (LRMs) have demonstrated impressive performance across a range of reasoning tasks, yet little is known about their internal reasoning processes in multilingual settings. We begin with a critical question: {\it In which language do these models reason when solving problems presented in different languages?} Our findings reveal that, despite multilingual training, LRMs tend to default to reasoning in high-resource languages (e.g., English) at test time, regardless of the input language. When constrained to reason in the same language as the input, model performance declines, especially for low-resource languages. In contrast, reasoning in high-resource languages generally preserves performance. We conduct extensive evaluations across reasoning-intensive tasks (MMMLU, MATH-500) and non-reasoning benchmarks (CulturalBench, LMSYS-toxic), showing that the effect of language choice varies by task type: input-language reasoning degrades performance on reasoning tasks but benefits cultural tasks, while safety evaluations exhibit language-specific behavior. By exposing these linguistic biases in LRMs, our work highlights a critical step toward developing more equitable models that serve users across diverse linguistic backgrounds. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17407 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? Tam, Zhi Rui Wu, Cheng-Kuang Chiu, Yu Ying Lin, Chieh-Yen Chen, Yun-Nung Lee, Hung-yi Computation and Language Large reasoning models (LRMs) have demonstrated impressive performance across a range of reasoning tasks, yet little is known about their internal reasoning processes in multilingual settings. We begin with a critical question: {\it In which language do these models reason when solving problems presented in different languages?} Our findings reveal that, despite multilingual training, LRMs tend to default to reasoning in high-resource languages (e.g., English) at test time, regardless of the input language. When constrained to reason in the same language as the input, model performance declines, especially for low-resource languages. In contrast, reasoning in high-resource languages generally preserves performance. We conduct extensive evaluations across reasoning-intensive tasks (MMMLU, MATH-500) and non-reasoning benchmarks (CulturalBench, LMSYS-toxic), showing that the effect of language choice varies by task type: input-language reasoning degrades performance on reasoning tasks but benefits cultural tasks, while safety evaluations exhibit language-specific behavior. By exposing these linguistic biases in LRMs, our work highlights a critical step toward developing more equitable models that serve users across diverse linguistic backgrounds. |
| title | Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.17407 |