NV-Retriever: Improving text embedding models with effective hard-negative mining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moreira, Gabriel de Souza P., Osmulski, Radek, Xu, Mengyao, Ak, Ronay, Schifferer, Benedikt, Oldridge, Even
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909481426747392
author Moreira, Gabriel de Souza P.
Osmulski, Radek
Xu, Mengyao
Ak, Ronay
Schifferer, Benedikt
Oldridge, Even
author_facet Moreira, Gabriel de Souza P.
Osmulski, Radek
Xu, Mengyao
Ak, Ronay
Schifferer, Benedikt
Oldridge, Even
contents Text embedding models have been popular for information retrieval applications such as semantic search and Question-Answering systems based on Retrieval-Augmented Generation (RAG). Those models are typically Transformer models that are fine-tuned with contrastive learning objectives. One of the challenging aspects of fine-tuning embedding models is the selection of high quality hard-negative passages for contrastive learning. In this paper we introduce a family of positive-aware mining methods that use the positive relevance score as an anchor for effective false negative removal, leading to faster training and more accurate retrieval models. We provide an ablation study on hard-negative mining methods over their configurations, exploring different teacher and base models. We further demonstrate the efficacy of our proposed mining methods at scale with the NV-Retriever-v1 model, which scores 60.9 on MTEB Retrieval (BEIR) benchmark and placed 1st when it was published to the MTEB Retrieval on July, 2024.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15831
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NV-Retriever: Improving text embedding models with effective hard-negative mining
Moreira, Gabriel de Souza P.
Osmulski, Radek
Xu, Mengyao
Ak, Ronay
Schifferer, Benedikt
Oldridge, Even
Information Retrieval
Artificial Intelligence
Text embedding models have been popular for information retrieval applications such as semantic search and Question-Answering systems based on Retrieval-Augmented Generation (RAG). Those models are typically Transformer models that are fine-tuned with contrastive learning objectives. One of the challenging aspects of fine-tuning embedding models is the selection of high quality hard-negative passages for contrastive learning. In this paper we introduce a family of positive-aware mining methods that use the positive relevance score as an anchor for effective false negative removal, leading to faster training and more accurate retrieval models. We provide an ablation study on hard-negative mining methods over their configurations, exploring different teacher and base models. We further demonstrate the efficacy of our proposed mining methods at scale with the NV-Retriever-v1 model, which scores 60.9 on MTEB Retrieval (BEIR) benchmark and placed 1st when it was published to the MTEB Retrieval on July, 2024.
title NV-Retriever: Improving text embedding models with effective hard-negative mining
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2407.15831