Under the Shadow of Babel: How Language Shapes Reasoning in LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Chenxi, Zhang, Yixuan, Gao, Lang, Xu, Zixiang, Song, Zirui, Wang, Yanbo, Chen, Xiuying
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913901475528704
author Wang, Chenxi
Zhang, Yixuan
Gao, Lang
Xu, Zixiang
Song, Zirui
Wang, Yanbo
Chen, Xiuying
author_facet Wang, Chenxi
Zhang, Yixuan
Gao, Lang
Xu, Zixiang
Song, Zirui
Wang, Yanbo
Chen, Xiuying
contents Language is not only a tool for communication but also a medium for human cognition and reasoning. If, as linguistic relativity suggests, the structure of language shapes cognitive patterns, then large language models (LLMs) trained on human language may also internalize the habitual logical structures embedded in different languages. To examine this hypothesis, we introduce BICAUSE, a structured bilingual dataset for causal reasoning, which includes semantically aligned Chinese and English samples in both forward and reversed causal forms. Our study reveals three key findings: (1) LLMs exhibit typologically aligned attention patterns, focusing more on causes and sentence-initial connectives in Chinese, while showing a more balanced distribution in English. (2) Models internalize language-specific preferences for causal word order and often rigidly apply them to atypical inputs, leading to degraded performance, especially in Chinese. (3) When causal reasoning succeeds, model representations converge toward semantically aligned abstractions across languages, indicating a shared understanding beyond surface form. Overall, these results suggest that LLMs not only mimic surface linguistic forms but also internalize the reasoning biases shaped by language. Rooted in cognitive linguistic theory, this phenomenon is for the first time empirically verified through structural analysis of model internals.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16151
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
Wang, Chenxi
Zhang, Yixuan
Gao, Lang
Xu, Zixiang
Song, Zirui
Wang, Yanbo
Chen, Xiuying
Computation and Language
Artificial Intelligence
Language is not only a tool for communication but also a medium for human cognition and reasoning. If, as linguistic relativity suggests, the structure of language shapes cognitive patterns, then large language models (LLMs) trained on human language may also internalize the habitual logical structures embedded in different languages. To examine this hypothesis, we introduce BICAUSE, a structured bilingual dataset for causal reasoning, which includes semantically aligned Chinese and English samples in both forward and reversed causal forms. Our study reveals three key findings: (1) LLMs exhibit typologically aligned attention patterns, focusing more on causes and sentence-initial connectives in Chinese, while showing a more balanced distribution in English. (2) Models internalize language-specific preferences for causal word order and often rigidly apply them to atypical inputs, leading to degraded performance, especially in Chinese. (3) When causal reasoning succeeds, model representations converge toward semantically aligned abstractions across languages, indicating a shared understanding beyond surface form. Overall, these results suggest that LLMs not only mimic surface linguistic forms but also internalize the reasoning biases shaped by language. Rooted in cognitive linguistic theory, this phenomenon is for the first time empirically verified through structural analysis of model internals.
title Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.16151