MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Banar, Nikolay, Lotfi, Ehsan, Van Nooten, Jens, Arhiliuc, Cristina, Kliocaite, Marija, Daelemans, Walter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914039061282816
author Banar, Nikolay
Lotfi, Ehsan
Van Nooten, Jens
Arhiliuc, Cristina
Kliocaite, Marija
Daelemans, Walter
author_facet Banar, Nikolay
Lotfi, Ehsan
Van Nooten, Jens
Arhiliuc, Cristina
Kliocaite, Marija
Daelemans, Walter
contents Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages. However, the Dutch language remains underrepresented, typically comprising only a small fraction of the published multilingual resources. To address this gap and encourage the further development of Dutch embeddings, we introduce new resources for their evaluation and generation. First, we introduce the Massive Text Embedding Benchmark for Dutch (MTEB-NL), which includes both existing Dutch datasets and newly created ones, covering a wide range of tasks. Second, we provide a training dataset compiled from available Dutch retrieval datasets, complemented with synthetic data generated by large language models to expand task coverage beyond retrieval. Finally, we release a series of E5-NL models compact yet efficient embedding models that demonstrate strong performance across multiple tasks. We make our resources publicly available through the Hugging Face Hub and the MTEB package.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
Banar, Nikolay
Lotfi, Ehsan
Van Nooten, Jens
Arhiliuc, Cristina
Kliocaite, Marija
Daelemans, Walter
Computation and Language
Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages. However, the Dutch language remains underrepresented, typically comprising only a small fraction of the published multilingual resources. To address this gap and encourage the further development of Dutch embeddings, we introduce new resources for their evaluation and generation. First, we introduce the Massive Text Embedding Benchmark for Dutch (MTEB-NL), which includes both existing Dutch datasets and newly created ones, covering a wide range of tasks. Second, we provide a training dataset compiled from available Dutch retrieval datasets, complemented with synthetic data generated by large language models to expand task coverage beyond retrieval. Finally, we release a series of E5-NL models compact yet efficient embedding models that demonstrate strong performance across multiple tasks. We make our resources publicly available through the Hugging Face Hub and the MTEB package.
title MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
topic Computation and Language
url https://arxiv.org/abs/2509.12340