Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bandarkar, Lucas, Ansell, Alan, Cohn, Trevor
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917351119650816
author Bandarkar, Lucas
Ansell, Alan
Cohn, Trevor
author_facet Bandarkar, Lucas
Ansell, Alan
Cohn, Trevor
contents Modern LLMs continue to exhibit significant variance in behavior across languages, such as being able to recall factual information in some languages but not others. While typically studied as a problem to be mitigated, in this work, we propose leveraging this cross-lingual inconsistency as a tool for interpretability in mixture-of-experts (MoE) LLMs. Our knowledge localization framework contrasts routing for sets of languages where the model correctly recalls information from languages where it fails. This allows us to isolate model components that play a functional role in answering about a piece of knowledge. Our method proceeds in two stages: (1) querying the model with difficult factual questions across a diverse set of languages to generate "success" and "failure" activation buckets and then (2) applying a statistical contrastive analysis to the MoE router logits to identify experts important for knowledge. To validate the necessity of this small number of experts for answering a knowledge question, we deactivate them and re-ask the question. We find that despite only deactivating about 20 out of 6000 experts, the model no longer answers correctly in over 40% of cases. Generally, this method provides a realistic and scalable knowledge localization approach to address increasingly complex LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
Bandarkar, Lucas
Ansell, Alan
Cohn, Trevor
Computation and Language
Artificial Intelligence
Machine Learning
Modern LLMs continue to exhibit significant variance in behavior across languages, such as being able to recall factual information in some languages but not others. While typically studied as a problem to be mitigated, in this work, we propose leveraging this cross-lingual inconsistency as a tool for interpretability in mixture-of-experts (MoE) LLMs. Our knowledge localization framework contrasts routing for sets of languages where the model correctly recalls information from languages where it fails. This allows us to isolate model components that play a functional role in answering about a piece of knowledge. Our method proceeds in two stages: (1) querying the model with difficult factual questions across a diverse set of languages to generate "success" and "failure" activation buckets and then (2) applying a statistical contrastive analysis to the MoE router logits to identify experts important for knowledge. To validate the necessity of this small number of experts for answering a knowledge question, we deactivate them and re-ask the question. We find that despite only deactivating about 20 out of 6000 experts, the model no longer answers correctly in over 40% of cases. Generally, this method provides a realistic and scalable knowledge localization approach to address increasingly complex LLMs.
title Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.17102