Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mukherjee, Anjishnu, Meng, Chutong, Anastasopoulos, Antonios
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918477934100480
author Mukherjee, Anjishnu
Meng, Chutong
Anastasopoulos, Antonios
author_facet Mukherjee, Anjishnu
Meng, Chutong
Anastasopoulos, Antonios
contents This paper argues that contemporary multilingual NLP has converged on a fragile and misleading paradigm of incidental multilingualism. Today's LLMs appear multilingual largely because they are trained on massive, uneven web corpora, not because multilingual or multicultural competence has been treated as a core design objective. We contend that this paradigm systematically produces unequal, brittle, and opaque behavior across languages, with severe consequences in real-world and agentic deployments where models must reason, plan, and act across multiple linguistic contexts. We report a focused empirical study of two practical questions: which languages models self-report as supported and which languages they actually respond in across multilingual prompts. We additionally demonstrate how even a simple language-change attack can surface these failures and expose hidden assumptions about language in LLM-based systems. To address this, we call for a shift toward multilingualism by design: a research agenda that treats equitable multilingual performance, cultural grounding, and cross-lingual behavioral understanding as first-class goals in all aspects of the model pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01224
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs
Mukherjee, Anjishnu
Meng, Chutong
Anastasopoulos, Antonios
Computation and Language
This paper argues that contemporary multilingual NLP has converged on a fragile and misleading paradigm of incidental multilingualism. Today's LLMs appear multilingual largely because they are trained on massive, uneven web corpora, not because multilingual or multicultural competence has been treated as a core design objective. We contend that this paradigm systematically produces unequal, brittle, and opaque behavior across languages, with severe consequences in real-world and agentic deployments where models must reason, plan, and act across multiple linguistic contexts. We report a focused empirical study of two practical questions: which languages models self-report as supported and which languages they actually respond in across multilingual prompts. We additionally demonstrate how even a simple language-change attack can surface these failures and expose hidden assumptions about language in LLM-based systems. To address this, we call for a shift toward multilingualism by design: a research agenda that treats equitable multilingual performance, cultural grounding, and cross-lingual behavioral understanding as first-class goals in all aspects of the model pipeline.
title Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs
topic Computation and Language
url https://arxiv.org/abs/2605.01224