ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ziyin, Liao, Zihan, Yu, Hang, Di, Peng, Wang, Rui
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910222281342976
author Zhang, Ziyin
Liao, Zihan
Yu, Hang
Di, Peng
Wang, Rui
author_facet Zhang, Ziyin
Liao, Zihan
Yu, Hang
Di, Peng
Wang, Rui
contents The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitive computational costs, a narrow linguistic focus that neglects most of the world's languages, and a lack of transparency from closed-source or open-weight models that stifles research. To dismantle these barriers, we introduce ML-Embed, a suite of inclusive and efficient models built upon a new framework: 3-Dimensional Matryoshka Learning (3D-ML). Our framework addresses the computational challenge with comprehensive efficiency across the entire model lifecycle. Beyond the storage benefits of Matryoshka Representation Learning (MRL) and flexible inference-time depth provided by Matryoshka Layer Learning (MLL), we introduce Matryoshka Embedding Learning (MEL) for enhanced parameter efficiency. To address the linguistic challenge, we curate a massively multilingual dataset and train a suite of models ranging from 140M to 8B parameters. In a direct commitment to transparency, we release all models, data, and code. Extensive evaluation on 430 tasks demonstrates that our models set new records on 9 of 17 evaluated MTEB benchmarks, with particularly strong results in low-resource languages, providing a reproducible blueprint for building globally equitable and computationally efficient AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15081
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Zhang, Ziyin
Liao, Zihan
Yu, Hang
Di, Peng
Wang, Rui
Computation and Language
Artificial Intelligence
The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitive computational costs, a narrow linguistic focus that neglects most of the world's languages, and a lack of transparency from closed-source or open-weight models that stifles research. To dismantle these barriers, we introduce ML-Embed, a suite of inclusive and efficient models built upon a new framework: 3-Dimensional Matryoshka Learning (3D-ML). Our framework addresses the computational challenge with comprehensive efficiency across the entire model lifecycle. Beyond the storage benefits of Matryoshka Representation Learning (MRL) and flexible inference-time depth provided by Matryoshka Layer Learning (MLL), we introduce Matryoshka Embedding Learning (MEL) for enhanced parameter efficiency. To address the linguistic challenge, we curate a massively multilingual dataset and train a suite of models ranging from 140M to 8B parameters. In a direct commitment to transparency, we release all models, data, and code. Extensive evaluation on 430 tasks demonstrates that our models set new records on 9 of 17 evaluated MTEB benchmarks, with particularly strong results in low-resource languages, providing a reproducible blueprint for building globally equitable and computationally efficient AI systems.
title ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.15081