The Wisdom of Many Queries: Complexity-Diversity Principle for Dense Retriever Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Xincan, Nishida, Noriki, Sakai, Yusuke, Matsumoto, Yuji
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915866542604288
author Feng, Xincan
Nishida, Noriki
Sakai, Yusuke
Matsumoto, Yuji
author_facet Feng, Xincan
Nishida, Noriki
Sakai, Yusuke
Matsumoto, Yuji
contents Synthetic query generation has become essential for training dense retrievers, yet prior methods generate one query per document, focusing solely on query quality. We are the first to systematically study multi-query synthesis and discover a quality-diversity trade-off: high-quality queries benefit in-domain tasks, while diverse queries benefit out-of-domain (OOD) generalization. Through controlled experiments on 4 benchmark types across Contriever, RetroMAE, and Qwen3-Embedding, we find that diversity benefit strongly correlates with query complexity (r$\geq$0.95, p<0.05), approximated by content words (CW). We formalize this as the Complexity-Diversity Principle (CDP): query complexity determines optimal diversity. Based on CDP, we propose complexity-aware training: multi-query synthesis for high-complexity tasks and CW-weighted training for existing data. Both strategies improve OOD performance on reasoning-intensive benchmarks, with compounded gains when combined.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09448
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Wisdom of Many Queries: Complexity-Diversity Principle for Dense Retriever Training
Feng, Xincan
Nishida, Noriki
Sakai, Yusuke
Matsumoto, Yuji
Information Retrieval
Machine Learning
Synthetic query generation has become essential for training dense retrievers, yet prior methods generate one query per document, focusing solely on query quality. We are the first to systematically study multi-query synthesis and discover a quality-diversity trade-off: high-quality queries benefit in-domain tasks, while diverse queries benefit out-of-domain (OOD) generalization. Through controlled experiments on 4 benchmark types across Contriever, RetroMAE, and Qwen3-Embedding, we find that diversity benefit strongly correlates with query complexity (r$\geq$0.95, p<0.05), approximated by content words (CW). We formalize this as the Complexity-Diversity Principle (CDP): query complexity determines optimal diversity. Based on CDP, we propose complexity-aware training: multi-query synthesis for high-complexity tasks and CW-weighted training for existing data. Both strategies improve OOD performance on reasoning-intensive benchmarks, with compounded gains when combined.
title The Wisdom of Many Queries: Complexity-Diversity Principle for Dense Retriever Training
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2602.09448