Augmenting Researchy Questions with Sub-question Judgments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ju, Jia-Huei, Yang, Eugene, Adriaanse, Trevor, Yates, Andrew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917041360863232
author Ju, Jia-Huei
Yang, Eugene
Adriaanse, Trevor
Yates, Andrew
author_facet Ju, Jia-Huei
Yang, Eugene
Adriaanse, Trevor
Yates, Andrew
contents The Researchy Questions dataset provides about 100k question queries with complex information needs that require retrieving information about several aspects of a topic. Each query in ResearchyQuestions is associated with sub-questions that were produced by prompting GPT-4. While ResearchyQuestions contains labels indicating what documents were clicked after issuing the query, there are no associations in the dataset between sub-questions and relevant documents. In this work, we augment the Researchy Questions dataset with LLM-judged labels for each sub-question using a Llama3.3 70B model. We intend these sub-question labels to serve as a resource for training retrieval models that better support complex information needs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21733
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Augmenting Researchy Questions with Sub-question Judgments
Ju, Jia-Huei
Yang, Eugene
Adriaanse, Trevor
Yates, Andrew
Information Retrieval
The Researchy Questions dataset provides about 100k question queries with complex information needs that require retrieving information about several aspects of a topic. Each query in ResearchyQuestions is associated with sub-questions that were produced by prompting GPT-4. While ResearchyQuestions contains labels indicating what documents were clicked after issuing the query, there are no associations in the dataset between sub-questions and relevant documents. In this work, we augment the Researchy Questions dataset with LLM-judged labels for each sub-question using a Llama3.3 70B model. We intend these sub-question labels to serve as a resource for training retrieval models that better support complex information needs.
title Augmenting Researchy Questions with Sub-question Judgments
topic Information Retrieval
url https://arxiv.org/abs/2510.21733