RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Saji, Alan, Husain, Jaavid Aktar, Jayakumar, Thanmay, Dabre, Raj, Kunchukuttan, Anoop, Puduppully, Ratish
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909972847132672
author Saji, Alan
Husain, Jaavid Aktar
Jayakumar, Thanmay
Dabre, Raj
Kunchukuttan, Anoop
Puduppully, Ratish
author_facet Saji, Alan
Husain, Jaavid Aktar
Jayakumar, Thanmay
Dabre, Raj
Kunchukuttan, Anoop
Puduppully, Ratish
contents Large Language Models (LLMs) exhibit strong multilingual performance despite being predominantly trained on English-centric corpora. This raises a fundamental question: How do LLMs achieve such multilingual capabilities? Focusing on languages written in non-Roman scripts, we investigate the role of Romanization - the representation of non-Roman scripts using Roman characters - as a potential bridge in multilingual processing. Using mechanistic interpretability techniques, we analyze next-token generation and find that intermediate layers frequently represent target words in Romanized form before transitioning to native script, a phenomenon we term Latent Romanization. Further, through activation patching experiments, we demonstrate that LLMs encode semantic concepts similarly across native and Romanized scripts, suggesting a shared underlying representation. Additionally, for translation into non-Roman script languages, our findings reveal that when the target language is in Romanized form, its representations emerge earlier in the model's layers compared to native script. These insights contribute to a deeper understanding of multilingual representation in LLMs and highlight the implicit role of Romanization in facilitating language transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07424
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
Saji, Alan
Husain, Jaavid Aktar
Jayakumar, Thanmay
Dabre, Raj
Kunchukuttan, Anoop
Puduppully, Ratish
Computation and Language
Artificial Intelligence
I.2.7
I.2.7
Large Language Models (LLMs) exhibit strong multilingual performance despite being predominantly trained on English-centric corpora. This raises a fundamental question: How do LLMs achieve such multilingual capabilities? Focusing on languages written in non-Roman scripts, we investigate the role of Romanization - the representation of non-Roman scripts using Roman characters - as a potential bridge in multilingual processing. Using mechanistic interpretability techniques, we analyze next-token generation and find that intermediate layers frequently represent target words in Romanized form before transitioning to native script, a phenomenon we term Latent Romanization. Further, through activation patching experiments, we demonstrate that LLMs encode semantic concepts similarly across native and Romanized scripts, suggesting a shared underlying representation. Additionally, for translation into non-Roman script languages, our findings reveal that when the target language is in Romanized form, its representations emerge earlier in the model's layers compared to native script. These insights contribute to a deeper understanding of multilingual representation in LLMs and highlight the implicit role of Romanization in facilitating language transfer.
title RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
topic Computation and Language
Artificial Intelligence
I.2.7
I.2.7
url https://arxiv.org/abs/2502.07424