MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ciancone, Mathieu, Kerboua, Imene, Schaeffer, Marion, Siblini, Wissam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914836467679232
author Ciancone, Mathieu
Kerboua, Imene
Schaeffer, Marion
Siblini, Wissam
author_facet Ciancone, Mathieu
Kerboua, Imene
Schaeffer, Marion
Siblini, Wissam
contents Recently, numerous embedding models have been made available and widely used for various NLP tasks. The Massive Text Embedding Benchmark (MTEB) has primarily simplified the process of choosing a model that performs well for several tasks in English, but extensions to other languages remain challenging. This is why we expand MTEB to propose the first massive benchmark of sentence embeddings for French. We gather 15 existing datasets in an easy-to-use interface and create three new French datasets for a global evaluation of 8 task categories. We compare 51 carefully selected embedding models on a large scale, conduct comprehensive statistical tests, and analyze the correlation between model performance and many of their characteristics. We find out that even if no model is the best on all tasks, large multilingual models pre-trained on sentence similarity perform exceptionally well. Our work comes with open-source code, new datasets and a public leaderboard.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20468
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
Ciancone, Mathieu
Kerboua, Imene
Schaeffer, Marion
Siblini, Wissam
Computation and Language
Information Retrieval
Machine Learning
Recently, numerous embedding models have been made available and widely used for various NLP tasks. The Massive Text Embedding Benchmark (MTEB) has primarily simplified the process of choosing a model that performs well for several tasks in English, but extensions to other languages remain challenging. This is why we expand MTEB to propose the first massive benchmark of sentence embeddings for French. We gather 15 existing datasets in an easy-to-use interface and create three new French datasets for a global evaluation of 8 task categories. We compare 51 carefully selected embedding models on a large scale, conduct comprehensive statistical tests, and analyze the correlation between model performance and many of their characteristics. We find out that even if no model is the best on all tasks, large multilingual models pre-trained on sentence similarity perform exceptionally well. Our work comes with open-source code, new datasets and a public leaderboard.
title MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
topic Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2405.20468