Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tamber, Manveer Singh, Kazi, Suleman, Sourabh, Vivek, Lin, Jimmy
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917938946113536
author Tamber, Manveer Singh
Kazi, Suleman
Sourabh, Vivek
Lin, Jimmy
author_facet Tamber, Manveer Singh
Kazi, Suleman
Sourabh, Vivek
Lin, Jimmy
contents While the current state-of-the-art dense retrieval models exhibit strong out-of-domain generalization, they might fail to capture nuanced domain-specific knowledge. In principle, fine-tuning these models for specialized retrieval tasks should yield higher effectiveness than relying on a one-size-fits-all model, but in practice, results can disappoint. We show that standard fine-tuning methods using an InfoNCE loss can unexpectedly degrade effectiveness rather than improve it, even for domain-specific scenarios. This holds true even when applying widely adopted techniques such as hard-negative mining and negative de-noising. To address this, we explore a training strategy that uses listwise distillation from a teacher cross-encoder, leveraging rich relevance signals to fine-tune the retriever. We further explore synthetic query generation using large language models. Through listwise distillation and training with a diverse set of queries ranging from natural user searches and factual claims to keyword-based queries, we achieve consistent effectiveness gains across multiple datasets. Our results also reveal that synthetic queries can rival human-written queries in training utility. However, we also identify limitations, particularly in the effectiveness of cross-encoder teachers as a bottleneck. We release our code and scripts to encourage further research.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
Tamber, Manveer Singh
Kazi, Suleman
Sourabh, Vivek
Lin, Jimmy
Information Retrieval
While the current state-of-the-art dense retrieval models exhibit strong out-of-domain generalization, they might fail to capture nuanced domain-specific knowledge. In principle, fine-tuning these models for specialized retrieval tasks should yield higher effectiveness than relying on a one-size-fits-all model, but in practice, results can disappoint. We show that standard fine-tuning methods using an InfoNCE loss can unexpectedly degrade effectiveness rather than improve it, even for domain-specific scenarios. This holds true even when applying widely adopted techniques such as hard-negative mining and negative de-noising. To address this, we explore a training strategy that uses listwise distillation from a teacher cross-encoder, leveraging rich relevance signals to fine-tune the retriever. We further explore synthetic query generation using large language models. Through listwise distillation and training with a diverse set of queries ranging from natural user searches and factual claims to keyword-based queries, we achieve consistent effectiveness gains across multiple datasets. Our results also reveal that synthetic queries can rival human-written queries in training utility. However, we also identify limitations, particularly in the effectiveness of cross-encoder teachers as a bottleneck. We release our code and scripts to encourage further research.
title Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
topic Information Retrieval
url https://arxiv.org/abs/2502.19712