ConvSDG: Session Data Generation for Conversational Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mo, Fengran, Yi, Bole, Mao, Kelong, Qu, Chen, Huang, Kaiyu, Nie, Jian-Yun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910370645409792
author Mo, Fengran
Yi, Bole
Mao, Kelong
Qu, Chen
Huang, Kaiyu
Nie, Jian-Yun
author_facet Mo, Fengran
Yi, Bole
Mao, Kelong
Qu, Chen
Huang, Kaiyu
Nie, Jian-Yun
contents Conversational search provides a more convenient interface for users to search by allowing multi-turn interaction with the search engine. However, the effectiveness of the conversational dense retrieval methods is limited by the scarcity of training data required for their fine-tuning. Thus, generating more training conversational sessions with relevant labels could potentially improve search performance. Based on the promising capabilities of large language models (LLMs) on text generation, we propose ConvSDG, a simple yet effective framework to explore the feasibility of boosting conversational search by using LLM for session data generation. Within this framework, we design dialogue/session-level and query-level data generation with unsupervised and semi-supervised learning, according to the availability of relevance judgments. The generated data are used to fine-tune the conversational dense retriever. Extensive experiments on four widely used datasets demonstrate the effectiveness and broad applicability of our ConvSDG framework compared with several strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11335
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ConvSDG: Session Data Generation for Conversational Search
Mo, Fengran
Yi, Bole
Mao, Kelong
Qu, Chen
Huang, Kaiyu
Nie, Jian-Yun
Information Retrieval
Computation and Language
Conversational search provides a more convenient interface for users to search by allowing multi-turn interaction with the search engine. However, the effectiveness of the conversational dense retrieval methods is limited by the scarcity of training data required for their fine-tuning. Thus, generating more training conversational sessions with relevant labels could potentially improve search performance. Based on the promising capabilities of large language models (LLMs) on text generation, we propose ConvSDG, a simple yet effective framework to explore the feasibility of boosting conversational search by using LLM for session data generation. Within this framework, we design dialogue/session-level and query-level data generation with unsupervised and semi-supervised learning, according to the availability of relevance judgments. The generated data are used to fine-tune the conversational dense retriever. Extensive experiments on four widely used datasets demonstrate the effectiveness and broad applicability of our ConvSDG framework compared with several strong baselines.
title ConvSDG: Session Data Generation for Conversational Search
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2403.11335