Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ying, Shuangshuang, Wang, Zheyu, Peng, Yunjian, Chen, Jin, Wu, Yuhao, Lin, Hongbin, He, Dingyu, Liu, Siyi, Yu, Gengchen, Piao, YinZhu, Wu, Yuchen, Gui, Xin, Peng, Zhongyuan, Li, Xin, Du, Xeron, Qin, Libo, Cao, YiXin, Zhang, Ge, Huang, Stephen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910005306851328
author Ying, Shuangshuang
Wang, Zheyu
Peng, Yunjian
Chen, Jin
Wu, Yuhao
Lin, Hongbin
He, Dingyu
Liu, Siyi
Yu, Gengchen
Piao, YinZhu
Wu, Yuchen
Gui, Xin
Peng, Zhongyuan
Li, Xin
Du, Xeron
Qin, Libo
Cao, YiXin
Zhang, Ge
Huang, Stephen
author_facet Ying, Shuangshuang
Wang, Zheyu
Peng, Yunjian
Chen, Jin
Wu, Yuhao
Lin, Hongbin
He, Dingyu
Liu, Siyi
Yu, Gengchen
Piao, YinZhu
Wu, Yuchen
Gui, Xin
Peng, Zhongyuan
Li, Xin
Du, Xeron
Qin, Libo
Cao, YiXin
Zhang, Ge
Huang, Stephen
contents Despite strong performance on existing benchmarks, it remains unclear whether large language models can reason over genuinely novel scientific information. Most evaluations score end-to-end RAG pipelines, where reasoning is confounded with retrieval and toolchain choices, and the signal is further contaminated by parametric memorization and open-web volatility. We introduce DeR2, a controlled deep-research sandbox that isolates document-grounded reasoning while preserving core difficulties of deep search: multi-step synthesis, denoising, and evidence-based conclusion making. DeR2 decouples evidence access from reasoning via four regimes--Instruction-only, Concepts (gold concepts without documents), Related-only (only relevant documents), and Full-set (relevant documents plus topically related distractors)--yielding interpretable regime gaps that operationalize retrieval loss vs. reasoning loss and enable fine-grained error attribution. To prevent parametric leakage, we apply a two-phase validation that requires parametric failure without evidence while ensuring oracle-concept solvability. To ensure reproducibility, each instance provides a frozen document library (drawn from 2023-2025 theoretical papers) with expert-annotated concepts and validated rationales. Experiments across a diverse set of state-of-the-art foundation models reveal substantial variation and significant headroom: some models exhibit mode-switch fragility, performing worse with the Full-set than with Instruction-only, while others show structural concept misuse, correctly naming concepts but failing to execute them as procedures.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21937
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
Ying, Shuangshuang
Wang, Zheyu
Peng, Yunjian
Chen, Jin
Wu, Yuhao
Lin, Hongbin
He, Dingyu
Liu, Siyi
Yu, Gengchen
Piao, YinZhu
Wu, Yuchen
Gui, Xin
Peng, Zhongyuan
Li, Xin
Du, Xeron
Qin, Libo
Cao, YiXin
Zhang, Ge
Huang, Stephen
Artificial Intelligence
Despite strong performance on existing benchmarks, it remains unclear whether large language models can reason over genuinely novel scientific information. Most evaluations score end-to-end RAG pipelines, where reasoning is confounded with retrieval and toolchain choices, and the signal is further contaminated by parametric memorization and open-web volatility. We introduce DeR2, a controlled deep-research sandbox that isolates document-grounded reasoning while preserving core difficulties of deep search: multi-step synthesis, denoising, and evidence-based conclusion making. DeR2 decouples evidence access from reasoning via four regimes--Instruction-only, Concepts (gold concepts without documents), Related-only (only relevant documents), and Full-set (relevant documents plus topically related distractors)--yielding interpretable regime gaps that operationalize retrieval loss vs. reasoning loss and enable fine-grained error attribution. To prevent parametric leakage, we apply a two-phase validation that requires parametric failure without evidence while ensuring oracle-concept solvability. To ensure reproducibility, each instance provides a frozen document library (drawn from 2023-2025 theoretical papers) with expert-annotated concepts and validated rationales. Experiments across a diverse set of state-of-the-art foundation models reveal substantial variation and significant headroom: some models exhibit mode-switch fragility, performing worse with the Full-set than with Instruction-only, while others show structural concept misuse, correctly naming concepts but failing to execute them as procedures.
title Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
topic Artificial Intelligence
url https://arxiv.org/abs/2601.21937