ConvSDG: Session Data Generation for Conversational Search
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910370645409792 |
|---|---|
| author | Mo, Fengran Yi, Bole Mao, Kelong Qu, Chen Huang, Kaiyu Nie, Jian-Yun |
| author_facet | Mo, Fengran Yi, Bole Mao, Kelong Qu, Chen Huang, Kaiyu Nie, Jian-Yun |
| contents | Conversational search provides a more convenient interface for users to search by allowing multi-turn interaction with the search engine. However, the effectiveness of the conversational dense retrieval methods is limited by the scarcity of training data required for their fine-tuning. Thus, generating more training conversational sessions with relevant labels could potentially improve search performance. Based on the promising capabilities of large language models (LLMs) on text generation, we propose ConvSDG, a simple yet effective framework to explore the feasibility of boosting conversational search by using LLM for session data generation. Within this framework, we design dialogue/session-level and query-level data generation with unsupervised and semi-supervised learning, according to the availability of relevance judgments. The generated data are used to fine-tune the conversational dense retriever. Extensive experiments on four widely used datasets demonstrate the effectiveness and broad applicability of our ConvSDG framework compared with several strong baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_11335 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ConvSDG: Session Data Generation for Conversational Search Mo, Fengran Yi, Bole Mao, Kelong Qu, Chen Huang, Kaiyu Nie, Jian-Yun Information Retrieval Computation and Language Conversational search provides a more convenient interface for users to search by allowing multi-turn interaction with the search engine. However, the effectiveness of the conversational dense retrieval methods is limited by the scarcity of training data required for their fine-tuning. Thus, generating more training conversational sessions with relevant labels could potentially improve search performance. Based on the promising capabilities of large language models (LLMs) on text generation, we propose ConvSDG, a simple yet effective framework to explore the feasibility of boosting conversational search by using LLM for session data generation. Within this framework, we design dialogue/session-level and query-level data generation with unsupervised and semi-supervised learning, according to the availability of relevance judgments. The generated data are used to fine-tune the conversational dense retriever. Extensive experiments on four widely used datasets demonstrate the effectiveness and broad applicability of our ConvSDG framework compared with several strong baselines. |
| title | ConvSDG: Session Data Generation for Conversational Search |
| topic | Information Retrieval Computation and Language |
| url | https://arxiv.org/abs/2403.11335 |