The Emergence of Abstract Thought in Large Language Models Beyond Any Language

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Yuxin, Zhao, Yiran, Zhang, Yang, Zhang, An, Kawaguchi, Kenji, Joty, Shafiq, Li, Junnan, Chua, Tat-Seng, Shieh, Michael Qizhe, Zhang, Wenxuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908404605255680
author Chen, Yuxin
Zhao, Yiran
Zhang, Yang
Zhang, An
Kawaguchi, Kenji
Joty, Shafiq
Li, Junnan
Chua, Tat-Seng
Shieh, Michael Qizhe
Zhang, Wenxuan
author_facet Chen, Yuxin
Zhao, Yiran
Zhang, Yang
Zhang, An
Kawaguchi, Kenji
Joty, Shafiq
Li, Junnan
Chua, Tat-Seng
Shieh, Michael Qizhe
Zhang, Wenxuan
contents As large language models (LLMs) continue to advance, their capacity to function effectively across a diverse range of languages has shown marked improvement. Preliminary studies observe that the hidden activations of LLMs often resemble English, even when responding to non-English prompts. This has led to the widespread assumption that LLMs may "think" in English. However, more recent results showing strong multilingual performance, even surpassing English performance on specific tasks in other languages, challenge this view. In this work, we find that LLMs progressively develop a core language-agnostic parameter space-a remarkably small subset of parameters whose deactivation results in significant performance degradation across all languages. This compact yet critical set of parameters underlies the model's ability to generalize beyond individual languages, supporting the emergence of abstract thought that is not tied to any specific linguistic system. Specifically, we identify language-related neurons-those are consistently activated during the processing of particular languages, and categorize them as either shared (active across multiple languages) or exclusive (specific to one). As LLMs undergo continued development over time, we observe a marked increase in both the proportion and functional importance of shared neurons, while exclusive neurons progressively diminish in influence. These shared neurons constitute the backbone of the core language-agnostic parameter space, supporting the emergence of abstract thought. Motivated by these insights, we propose neuron-specific training strategies tailored to LLMs' language-agnostic levels at different development stages. Experiments across diverse LLM families support our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09890
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Emergence of Abstract Thought in Large Language Models Beyond Any Language
Chen, Yuxin
Zhao, Yiran
Zhang, Yang
Zhang, An
Kawaguchi, Kenji
Joty, Shafiq
Li, Junnan
Chua, Tat-Seng
Shieh, Michael Qizhe
Zhang, Wenxuan
Computation and Language
Artificial Intelligence
As large language models (LLMs) continue to advance, their capacity to function effectively across a diverse range of languages has shown marked improvement. Preliminary studies observe that the hidden activations of LLMs often resemble English, even when responding to non-English prompts. This has led to the widespread assumption that LLMs may "think" in English. However, more recent results showing strong multilingual performance, even surpassing English performance on specific tasks in other languages, challenge this view. In this work, we find that LLMs progressively develop a core language-agnostic parameter space-a remarkably small subset of parameters whose deactivation results in significant performance degradation across all languages. This compact yet critical set of parameters underlies the model's ability to generalize beyond individual languages, supporting the emergence of abstract thought that is not tied to any specific linguistic system. Specifically, we identify language-related neurons-those are consistently activated during the processing of particular languages, and categorize them as either shared (active across multiple languages) or exclusive (specific to one). As LLMs undergo continued development over time, we observe a marked increase in both the proportion and functional importance of shared neurons, while exclusive neurons progressively diminish in influence. These shared neurons constitute the backbone of the core language-agnostic parameter space, supporting the emergence of abstract thought. Motivated by these insights, we propose neuron-specific training strategies tailored to LLMs' language-agnostic levels at different development stages. Experiments across diverse LLM families support our approach.
title The Emergence of Abstract Thought in Large Language Models Beyond Any Language
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.09890