The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
Fuente:
arXiv
Saved in:
| Main Authors: | Shani, Chen, Reif, Yuval, Roll, Nathan, Jurafsky, Dan, Shutova, Ekaterina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
by: Zhang, Christine, et al.
Published: (2026)
by: Zhang, Christine, et al.
Published: (2026)
PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation
by: Roll, Nathan
Published: (2025)
by: Roll, Nathan
Published: (2025)
Entanglement as Memory: Mechanistic Interpretability of Quantum Language Models
by: Roll, Nathan
Published: (2026)
by: Roll, Nathan
Published: (2026)
False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
by: Kallini, Julie, et al.
Published: (2025)
by: Kallini, Julie, et al.
Published: (2025)
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
by: Choenni, Rochelle, et al.
Published: (2023)
by: Choenni, Rochelle, et al.
Published: (2023)
Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition
by: Roll, Nathan, et al.
Published: (2026)
by: Roll, Nathan, et al.
Published: (2026)
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs
by: Reif, Yuval, et al.
Published: (2024)
by: Reif, Yuval, et al.
Published: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
by: Roll, Nathan, et al.
Published: (2025)
by: Roll, Nathan, et al.
Published: (2025)
Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning
by: Crosbie, Joy, et al.
Published: (2024)
by: Crosbie, Joy, et al.
Published: (2024)
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
Beyond Tokens: Concept-Level Training Objectives for LLMs
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
Cross-modal Information Flow in Multimodal Large Language Models
by: Zhang, Zhi, et al.
Published: (2024)
by: Zhang, Zhi, et al.
Published: (2024)
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
by: Verhoeven, Ivo, et al.
Published: (2024)
by: Verhoeven, Ivo, et al.
Published: (2024)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
by: Zhou, Kaitlyn, et al.
Published: (2025)
by: Zhou, Kaitlyn, et al.
Published: (2025)
Quantifying Language Disparities in Multilingual Large Language Models
by: Hu, Songbo, et al.
Published: (2025)
by: Hu, Songbo, et al.
Published: (2025)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
by: Mizrahi, Moran, et al.
Published: (2025)
by: Mizrahi, Moran, et al.
Published: (2025)
Artificial Aphasias in Lesioned Language Models
by: Roll, Nathan, et al.
Published: (2026)
by: Roll, Nathan, et al.
Published: (2026)
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
by: Rajaee, Sara, et al.
Published: (2025)
by: Rajaee, Sara, et al.
Published: (2025)
A framework for annotating and modelling intentions behind metaphor use
by: Michelli, Gianluca, et al.
Published: (2024)
by: Michelli, Gianluca, et al.
Published: (2024)
Density Matrices for Metaphor Understanding
by: Owers, Jay, et al.
Published: (2024)
by: Owers, Jay, et al.
Published: (2024)
A Shared Geometry of Difficulty in Multilingual Language Models
by: Civelli, Stefano, et al.
Published: (2026)
by: Civelli, Stefano, et al.
Published: (2026)
Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
by: Ògúnrèmí, Tolúlopé, et al.
Published: (2025)
by: Ògúnrèmí, Tolúlopé, et al.
Published: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
by: Bartelds, Martijn, et al.
Published: (2025)
by: Bartelds, Martijn, et al.
Published: (2025)
Metaphor Understanding Challenge Dataset for LLMs
by: Tong, Xiaoyu, et al.
Published: (2024)
by: Tong, Xiaoyu, et al.
Published: (2024)
Grounding Gaps in Language Model Generations
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
by: Leidinger, Alina, et al.
Published: (2024)
by: Leidinger, Alina, et al.
Published: (2024)
Learning New Tasks from a Few Examples with Soft-Label Prototypes
by: Singh, Avyav Kumar, et al.
Published: (2022)
by: Singh, Avyav Kumar, et al.
Published: (2022)
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
by: Reif, Yuval, et al.
Published: (2025)
by: Reif, Yuval, et al.
Published: (2025)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
by: Kaplan, Guy, et al.
Published: (2025)
by: Kaplan, Guy, et al.
Published: (2025)
Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
by: Tomikawa, Yuto, et al.
Published: (2025)
by: Tomikawa, Yuto, et al.
Published: (2025)
Exploring Representational Disparities Between Multilingual and Bilingual Translation Models
by: Verma, Neha, et al.
Published: (2023)
by: Verma, Neha, et al.
Published: (2023)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
by: Verhoeven, Ivo, et al.
Published: (2024)
by: Verhoeven, Ivo, et al.
Published: (2024)
The Script Tax: Measuring Tokenization-Driven Efficiency and Latency Disparities in Multilingual Language Models
by: Dixit, Aradhya, et al.
Published: (2026)
by: Dixit, Aradhya, et al.
Published: (2026)
Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?
by: Zotos, Leonidas, et al.
Published: (2024)
by: Zotos, Leonidas, et al.
Published: (2024)
A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models
by: de la Fuente, Antón, et al.
Published: (2024)
by: de la Fuente, Antón, et al.
Published: (2024)
Similar Items
-
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
by: Zhang, Christine, et al.
Published: (2026) -
PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation
by: Roll, Nathan
Published: (2025) -
Entanglement as Memory: Mechanistic Interpretability of Quantum Language Models
by: Roll, Nathan
Published: (2026) -
False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
by: Kallini, Julie, et al.
Published: (2025) -
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
by: Choenni, Rochelle, et al.
Published: (2024)