Negative Sampling Techniques in Information Retrieval: A Survey

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wischounig, Laurin, Abdallah, Abdelrahman, Jatowt, Adam
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917352256307200
author Wischounig, Laurin
Abdallah, Abdelrahman
Jatowt, Adam
author_facet Wischounig, Laurin
Abdallah, Abdelrahman
Jatowt, Adam
contents Information Retrieval (IR) is fundamental to many modern NLP applications. The rise of dense retrieval (DR), using neural networks to learn semantic vector representations, has significantly advanced IR performance. Central to training effective dense retrievers through contrastive learning is the selection of informative negative samples. Synthesizing 35 seminal papers, this survey provides a comprehensive and up-to-date overview of negative sampling techniques in dense IR. Our unique contribution is the focus on modern NLP applications and the inclusion of recent Large Language Model (LLM)-driven methods, an area absent in prior reviews. We propose a taxonomy that categorizes techniques including random, static/dynamically mined, and synthetic datasets. We then analyze these approaches with respect to trade-offs between effectiveness, computational cost, and implementation difficulty. The survey concludes by outlining current challenges and promising future directions for the use of LLM-generated synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18005
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Negative Sampling Techniques in Information Retrieval: A Survey
Wischounig, Laurin
Abdallah, Abdelrahman
Jatowt, Adam
Information Retrieval
Information Retrieval (IR) is fundamental to many modern NLP applications. The rise of dense retrieval (DR), using neural networks to learn semantic vector representations, has significantly advanced IR performance. Central to training effective dense retrievers through contrastive learning is the selection of informative negative samples. Synthesizing 35 seminal papers, this survey provides a comprehensive and up-to-date overview of negative sampling techniques in dense IR. Our unique contribution is the focus on modern NLP applications and the inclusion of recent Large Language Model (LLM)-driven methods, an area absent in prior reviews. We propose a taxonomy that categorizes techniques including random, static/dynamically mined, and synthetic datasets. We then analyze these approaches with respect to trade-offs between effectiveness, computational cost, and implementation difficulty. The survey concludes by outlining current challenges and promising future directions for the use of LLM-generated synthetic data.
title Negative Sampling Techniques in Information Retrieval: A Survey
topic Information Retrieval
url https://arxiv.org/abs/2603.18005