Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Merrick, Luke, Xu, Danmei, Nuti, Gaurav, Campos, Daniel
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910439380615168
author Merrick, Luke
Xu, Danmei
Nuti, Gaurav
Campos, Daniel
author_facet Merrick, Luke
Xu, Danmei
Nuti, Gaurav
Campos, Daniel
contents This report describes the training dataset creation and recipe behind the family of \texttt{arctic-embed} text embedding models (a set of five models ranging from 22 to 334 million parameters with weights open-sourced under an Apache-2 license). At the time of their release, each model achieved state-of-the-art retrieval accuracy for models of their size on the MTEB Retrieval leaderboard, with the largest model, arctic-embed-l outperforming closed source embedding models such as Cohere's embed-v3 and Open AI's text-embed-3-large. In addition to the details of our training recipe, we have provided several informative ablation studies, which we believe are the cause of our model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05374
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
Merrick, Luke
Xu, Danmei
Nuti, Gaurav
Campos, Daniel
Computation and Language
Artificial Intelligence
Information Retrieval
This report describes the training dataset creation and recipe behind the family of \texttt{arctic-embed} text embedding models (a set of five models ranging from 22 to 334 million parameters with weights open-sourced under an Apache-2 license). At the time of their release, each model achieved state-of-the-art retrieval accuracy for models of their size on the MTEB Retrieval leaderboard, with the largest model, arctic-embed-l outperforming closed source embedding models such as Cohere's embed-v3 and Open AI's text-embed-3-large. In addition to the details of our training recipe, we have provided several informative ablation studies, which we believe are the cause of our model performance.
title Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2405.05374