ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mo, Fengran, Zhang, Jinghan, Hui, Yuchen, Sun, Jia Ao, Xu, Zhichao, Su, Zhan, Nie, Jian-Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908646755008512
author Mo, Fengran
Zhang, Jinghan
Hui, Yuchen
Sun, Jia Ao
Xu, Zhichao
Su, Zhan
Nie, Jian-Yun
author_facet Mo, Fengran
Zhang, Jinghan
Hui, Yuchen
Sun, Jia Ao
Xu, Zhichao
Su, Zhan
Nie, Jian-Yun
contents Conversational search aims to satisfy users' complex information needs via multiple-turn interactions. The key challenge lies in revealing real users' search intent from the context-dependent queries. Previous studies achieve conversational search by fine-tuning a conversational dense retriever with relevance judgments between pairs of context-dependent queries and documents. However, this training paradigm encounters data scarcity issues. To this end, we propose ConvMix, a mixed-criteria framework to augment conversational dense retrieval, which covers more aspects than existing data augmentation frameworks. We design a two-sided relevance judgment augmentation schema in a scalable manner via the aid of large language models. Besides, we integrate the framework with quality control mechanisms to obtain semantically diverse samples and near-distribution supervisions to combine various annotated data. Experimental results on five widely used benchmarks show that the conversational dense retriever trained by our ConvMix framework outperforms previous baseline methods, which demonstrates our superior effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense Retrieval
Mo, Fengran
Zhang, Jinghan
Hui, Yuchen
Sun, Jia Ao
Xu, Zhichao
Su, Zhan
Nie, Jian-Yun
Information Retrieval
Computation and Language
Conversational search aims to satisfy users' complex information needs via multiple-turn interactions. The key challenge lies in revealing real users' search intent from the context-dependent queries. Previous studies achieve conversational search by fine-tuning a conversational dense retriever with relevance judgments between pairs of context-dependent queries and documents. However, this training paradigm encounters data scarcity issues. To this end, we propose ConvMix, a mixed-criteria framework to augment conversational dense retrieval, which covers more aspects than existing data augmentation frameworks. We design a two-sided relevance judgment augmentation schema in a scalable manner via the aid of large language models. Besides, we integrate the framework with quality control mechanisms to obtain semantically diverse samples and near-distribution supervisions to combine various annotated data. Experimental results on five widely used benchmarks show that the conversational dense retriever trained by our ConvMix framework outperforms previous baseline methods, which demonstrates our superior effectiveness.
title ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense Retrieval
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2508.04001