Language-Specific Latent Process Hinders Cross-Lingual Performance

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lim, Zheng Wei, Aji, Alham Fikri, Cohn, Trevor
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908559811280896
author Lim, Zheng Wei
Aji, Alham Fikri
Cohn, Trevor
author_facet Lim, Zheng Wei
Aji, Alham Fikri
Cohn, Trevor
contents Large language models (LLMs) are demonstrably capable of cross-lingual transfer, but can produce inconsistent output when prompted with the same queries written in different languages. To understand how language models are able to generalize knowledge from one language to the others, we measure representation similarity between languages, and apply the logit lens to interpret the implicit steps taken by LLMs to solve multilingual multi-choice reasoning questions. Our analyses reveal LLMs predict inconsistently and are less accurate because they rely on representations that are dissimilar across languages, rather than working in a shared semantic space. While larger models are more multilingual, we show their hidden states are more likely to dissociate from the shared representation compared to smaller models, but are nevertheless more capable of retrieving knowledge embedded across different languages. Finally, we demonstrate that knowledge sharing in small models can be facilitated by steering their latent processing towards the shared semantic space. This improves the models' multilingual reasoning performance, as a result of more knowledge transfer from, and better output consistency with English.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13141
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language-Specific Latent Process Hinders Cross-Lingual Performance
Lim, Zheng Wei
Aji, Alham Fikri
Cohn, Trevor
Computation and Language
Large language models (LLMs) are demonstrably capable of cross-lingual transfer, but can produce inconsistent output when prompted with the same queries written in different languages. To understand how language models are able to generalize knowledge from one language to the others, we measure representation similarity between languages, and apply the logit lens to interpret the implicit steps taken by LLMs to solve multilingual multi-choice reasoning questions. Our analyses reveal LLMs predict inconsistently and are less accurate because they rely on representations that are dissimilar across languages, rather than working in a shared semantic space. While larger models are more multilingual, we show their hidden states are more likely to dissociate from the shared representation compared to smaller models, but are nevertheless more capable of retrieving knowledge embedded across different languages. Finally, we demonstrate that knowledge sharing in small models can be facilitated by steering their latent processing towards the shared semantic space. This improves the models' multilingual reasoning performance, as a result of more knowledge transfer from, and better output consistency with English.
title Language-Specific Latent Process Hinders Cross-Lingual Performance
topic Computation and Language
url https://arxiv.org/abs/2505.13141