DEPART: DEcomposing PARiTy across Multilingual LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Uppadhyay, Manan, Kodali, Prashant, Chitale, Pranjal, Ramaprasad, Reshma, Beniwal, Himanshu, Sitaram, Sunayana
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918526756847616
author Uppadhyay, Manan
Kodali, Prashant
Chitale, Pranjal
Ramaprasad, Reshma
Beniwal, Himanshu
Sitaram, Sunayana
author_facet Uppadhyay, Manan
Kodali, Prashant
Chitale, Pranjal
Ramaprasad, Reshma
Beniwal, Himanshu
Sitaram, Sunayana
contents Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering practitioners no actionable levers. We first establish that these gaps are systematic rather than artifacts of sampling noise via distribution-free Friedman and Kruskal--Wallis tests, then introduce a two-step Bayesian hierarchical framework that decomposes multilingual performance variance into interpretable components. First, isolating the variance attributable to language identity, we show that observable language features (script, family, typological distance) explain $R^2_{\text{ling}} = 79\%$ of this variance on understanding tasks and $92\%$ on reasoning, with a model's internal representational similarity to English emerging as the dominant predictor across both task buckets. Second, decomposing the full (model$\times$benchmark$\times$language) cube, we find that NLU and reasoning have fundamentally divergent variance profiles: model identity dominates understanding ($66.7\%$ of variance), whereas the benchmark$\times$model interaction dominates reasoning ($46.3\%$). Together these results recast multilingual evaluation from passive performance mapping into an explainable, diagnostic framework with concrete levers for targeting the root drivers of language disparity.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28163
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DEPART: DEcomposing PARiTy across Multilingual LLMs
Uppadhyay, Manan
Kodali, Prashant
Chitale, Pranjal
Ramaprasad, Reshma
Beniwal, Himanshu
Sitaram, Sunayana
Computation and Language
Artificial Intelligence
Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering practitioners no actionable levers. We first establish that these gaps are systematic rather than artifacts of sampling noise via distribution-free Friedman and Kruskal--Wallis tests, then introduce a two-step Bayesian hierarchical framework that decomposes multilingual performance variance into interpretable components. First, isolating the variance attributable to language identity, we show that observable language features (script, family, typological distance) explain $R^2_{\text{ling}} = 79\%$ of this variance on understanding tasks and $92\%$ on reasoning, with a model's internal representational similarity to English emerging as the dominant predictor across both task buckets. Second, decomposing the full (model$\times$benchmark$\times$language) cube, we find that NLU and reasoning have fundamentally divergent variance profiles: model identity dominates understanding ($66.7\%$ of variance), whereas the benchmark$\times$model interaction dominates reasoning ($46.3\%$). Together these results recast multilingual evaluation from passive performance mapping into an explainable, diagnostic framework with concrete levers for targeting the root drivers of language disparity.
title DEPART: DEcomposing PARiTy across Multilingual LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.28163