Remining Hard Negatives for Generative Pseudo Labeled Domain Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuksel, Goksenin, Rau, David, Kamps, Jaap
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915120067641344
author Yuksel, Goksenin
Rau, David
Kamps, Jaap
author_facet Yuksel, Goksenin
Rau, David
Kamps, Jaap
contents Dense retrievers have demonstrated significant potential for neural information retrieval; however, they exhibit a lack of robustness to domain shifts, thereby limiting their efficacy in zero-shot settings across diverse domains. A state-of-the-art domain adaptation technique is Generative Pseudo Labeling (GPL). GPL uses synthetic query generation and initially mined hard negatives to distill knowledge from cross-encoder to dense retrievers in the target domain. In this paper, we analyze the documents retrieved by the domain-adapted model and discover that these are more relevant to the target queries than those of the non-domain-adapted model. We then propose refreshing the hard-negative index during the knowledge distillation phase to mine better hard negatives. Our remining R-GPL approach boosts ranking performance in 13/14 BEIR datasets and 9/12 LoTTe datasets. Our contributions are (i) analyzing hard negatives returned by domain-adapted and non-domain-adapted models and (ii) applying the GPL training with and without hard-negative re-mining in LoTTE and BEIR datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Remining Hard Negatives for Generative Pseudo Labeled Domain Adaptation
Yuksel, Goksenin
Rau, David
Kamps, Jaap
Information Retrieval
Machine Learning
Dense retrievers have demonstrated significant potential for neural information retrieval; however, they exhibit a lack of robustness to domain shifts, thereby limiting their efficacy in zero-shot settings across diverse domains. A state-of-the-art domain adaptation technique is Generative Pseudo Labeling (GPL). GPL uses synthetic query generation and initially mined hard negatives to distill knowledge from cross-encoder to dense retrievers in the target domain. In this paper, we analyze the documents retrieved by the domain-adapted model and discover that these are more relevant to the target queries than those of the non-domain-adapted model. We then propose refreshing the hard-negative index during the knowledge distillation phase to mine better hard negatives. Our remining R-GPL approach boosts ranking performance in 13/14 BEIR datasets and 9/12 LoTTe datasets. Our contributions are (i) analyzing hard negatives returned by domain-adapted and non-domain-adapted models and (ii) applying the GPL training with and without hard-negative re-mining in LoTTE and BEIR datasets.
title Remining Hard Negatives for Generative Pseudo Labeled Domain Adaptation
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2501.14434