On Synthetic Data Strategies for Domain-Specific Generative Retrieval

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen, Haoyang, Guo, Jiang, Zhang, Yi, Jiang, Jiarong, Wang, Zhiguo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916629067071488
author Wen, Haoyang
Guo, Jiang
Zhang, Yi
Jiang, Jiarong
Wang, Zhiguo
author_facet Wen, Haoyang
Guo, Jiang
Zhang, Yi
Jiang, Jiarong
Wang, Zhiguo
contents This paper investigates synthetic data generation strategies in developing generative retrieval models for domain-specific corpora, thereby addressing the scalability challenges inherent in manually annotating in-domain queries. We study the data strategies for a two-stage training framework: in the first stage, which focuses on learning to decode document identifiers from queries, we investigate LLM-generated queries across multiple granularity (e.g. chunks, sentences) and domain-relevant search constraints that can better capture nuanced relevancy signals. In the second stage, which aims to refine document ranking through preference learning, we explore the strategies for mining hard negatives based on the initial model's predictions. Experiments on public datasets over diverse domains demonstrate the effectiveness of our synthetic data generation and hard negative sampling approach.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17957
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Synthetic Data Strategies for Domain-Specific Generative Retrieval
Wen, Haoyang
Guo, Jiang
Zhang, Yi
Jiang, Jiarong
Wang, Zhiguo
Computation and Language
Information Retrieval
This paper investigates synthetic data generation strategies in developing generative retrieval models for domain-specific corpora, thereby addressing the scalability challenges inherent in manually annotating in-domain queries. We study the data strategies for a two-stage training framework: in the first stage, which focuses on learning to decode document identifiers from queries, we investigate LLM-generated queries across multiple granularity (e.g. chunks, sentences) and domain-relevant search constraints that can better capture nuanced relevancy signals. In the second stage, which aims to refine document ranking through preference learning, we explore the strategies for mining hard negatives based on the initial model's predictions. Experiments on public datasets over diverse domains demonstrate the effectiveness of our synthetic data generation and hard negative sampling approach.
title On Synthetic Data Strategies for Domain-Specific Generative Retrieval
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2502.17957