False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kallini, Julie, Jurafsky, Dan, Potts, Christopher, Bartelds, Martijn
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908556950765568
author Kallini, Julie
Jurafsky, Dan
Potts, Christopher
Bartelds, Martijn
author_facet Kallini, Julie
Jurafsky, Dan
Potts, Christopher
Bartelds, Martijn
contents Subword tokenizers trained on multilingual corpora naturally produce overlapping tokens across languages. Does token overlap facilitate cross-lingual transfer or instead introduce interference between languages? Prior work offers mixed evidence, partly due to varied setups and confounders, such as token frequency or subword segmentation granularity. To address this question, we devise a controlled experiment where we train bilingual autoregressive models on multiple language pairs under systematically varied vocabulary overlap settings. Crucially, we explore a new dimension to understanding how overlap affects transfer: the semantic similarity of tokens shared across languages. We first analyze our models' hidden representations and find that overlap of any kind creates embedding spaces that capture cross-lingual semantic relationships, while this effect is much weaker in models with disjoint vocabularies. On XNLI and XQuAD, we find that models with overlap outperform models with disjoint vocabularies, and that transfer performance generally improves as overlap increases. Overall, our findings highlight the advantages of token overlap in multilingual models and show that substantial shared vocabulary remains a beneficial design choice for multilingual tokenizers.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18750
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
Kallini, Julie
Jurafsky, Dan
Potts, Christopher
Bartelds, Martijn
Computation and Language
Subword tokenizers trained on multilingual corpora naturally produce overlapping tokens across languages. Does token overlap facilitate cross-lingual transfer or instead introduce interference between languages? Prior work offers mixed evidence, partly due to varied setups and confounders, such as token frequency or subword segmentation granularity. To address this question, we devise a controlled experiment where we train bilingual autoregressive models on multiple language pairs under systematically varied vocabulary overlap settings. Crucially, we explore a new dimension to understanding how overlap affects transfer: the semantic similarity of tokens shared across languages. We first analyze our models' hidden representations and find that overlap of any kind creates embedding spaces that capture cross-lingual semantic relationships, while this effect is much weaker in models with disjoint vocabularies. On XNLI and XQuAD, we find that models with overlap outperform models with disjoint vocabularies, and that transfer performance generally improves as overlap increases. Overall, our findings highlight the advantages of token overlap in multilingual models and show that substantial shared vocabulary remains a beneficial design choice for multilingual tokenizers.
title False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.18750