ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Junyong, Dai, Lu, Han, Ruiqian, Sui, Yijie, Wang, Ruilin, Sun, Xingliang, Wu, Qinglin, Feng, Min, Liu, Hao, Xiong, Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918057358655488
author Lin, Junyong
Dai, Lu
Han, Ruiqian
Sui, Yijie
Wang, Ruilin
Sun, Xingliang
Wu, Qinglin
Feng, Min
Liu, Hao
Xiong, Hui
author_facet Lin, Junyong
Dai, Lu
Han, Ruiqian
Sui, Yijie
Wang, Ruilin
Sun, Xingliang
Wu, Qinglin
Feng, Min
Liu, Hao
Xiong, Hui
contents Scientific researchers need intensive information about datasets to effectively evaluate and develop theories and methodologies. The information needs regarding datasets are implicitly embedded in particular research tasks, rather than explicitly expressed in search queries. However, existing scientific retrieval and question-answering (QA) datasets typically address straightforward questions, which do not align with the distribution of real-world research inquiries. To bridge this gap, we developed ScIRGen, a dataset generation framework for scientific QA \& retrieval that more accurately reflects the information needs of professional science researchers, and uses it to create a large-scale scientific retrieval-augmented generation (RAG) dataset with realistic queries, datasets and papers. Technically, we designed a dataset-oriented information extraction method that leverages academic papers to augment the dataset representation. We then proposed a question generation framework by employing cognitive taxonomy to ensure the quality of synthesized questions. We also design a method to automatically filter synthetic answers based on the perplexity shift of LLMs, which is highly aligned with human judgment of answers' validity. Collectively, these methodologies culminated in the creation of the 61k QA dataset, ScIRGen-Geo. We benchmarked representative methods on the ScIRGen-Geo dataset for their question-answering and retrieval capabilities, finding out that current methods still suffer from reasoning from complex questions. This work advances the development of more sophisticated tools to support the intricate information needs of the scientific community.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research
Lin, Junyong
Dai, Lu
Han, Ruiqian
Sui, Yijie
Wang, Ruilin
Sun, Xingliang
Wu, Qinglin
Feng, Min
Liu, Hao
Xiong, Hui
Computation and Language
Artificial Intelligence
Information Retrieval
Scientific researchers need intensive information about datasets to effectively evaluate and develop theories and methodologies. The information needs regarding datasets are implicitly embedded in particular research tasks, rather than explicitly expressed in search queries. However, existing scientific retrieval and question-answering (QA) datasets typically address straightforward questions, which do not align with the distribution of real-world research inquiries. To bridge this gap, we developed ScIRGen, a dataset generation framework for scientific QA \& retrieval that more accurately reflects the information needs of professional science researchers, and uses it to create a large-scale scientific retrieval-augmented generation (RAG) dataset with realistic queries, datasets and papers. Technically, we designed a dataset-oriented information extraction method that leverages academic papers to augment the dataset representation. We then proposed a question generation framework by employing cognitive taxonomy to ensure the quality of synthesized questions. We also design a method to automatically filter synthetic answers based on the perplexity shift of LLMs, which is highly aligned with human judgment of answers' validity. Collectively, these methodologies culminated in the creation of the 61k QA dataset, ScIRGen-Geo. We benchmarked representative methods on the ScIRGen-Geo dataset for their question-answering and retrieval capabilities, finding out that current methods still suffer from reasoning from complex questions. This work advances the development of more sophisticated tools to support the intricate information needs of the scientific community.
title ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2506.11117