DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
Fuente:
arXiv
Saved in:
| Main Authors: | Coelho, João, Ning, Jingjie, He, Jingyuan, Mao, Kangrui, Paladugu, Abhijay, Setlur, Pranav, Jin, Jiahe, Callan, Jamie, Magalhães, João, Martins, Bruno, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
by: Jin, Jiahe, et al.
Published: (2025)
by: Jin, Jiahe, et al.
Published: (2025)
Dwell in the Beginning: How Language Models Embed Long Documents for Dense Retrieval
by: Coelho, João, et al.
Published: (2024)
by: Coelho, João, et al.
Published: (2024)
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
Aligning Web Query Generation with Ranking Objectives via Direct Preference Optimization
by: Coelho, João, et al.
Published: (2025)
by: Coelho, João, et al.
Published: (2025)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests
by: He, Jingyuan, et al.
Published: (2025)
by: He, Jingyuan, et al.
Published: (2025)
A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis
by: Martins, Bruno, et al.
Published: (2025)
by: Martins, Bruno, et al.
Published: (2025)
Less LLM, More Documents: Searching for Improved RAG
by: Ning, Jingjie, et al.
Published: (2025)
by: Ning, Jingjie, et al.
Published: (2025)
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
by: Chandrahasan, Prahaladh, et al.
Published: (2025)
by: Chandrahasan, Prahaladh, et al.
Published: (2025)
Advancing Research Transparency and Reproducibility in Pharmacoepidemiology
by: Shirley V. Wang, et al.
Published: (2025)
by: Shirley V. Wang, et al.
Published: (2025)
Lisbon Computational Linguists at SemEval-2024 Task 2: Using A Mistral 7B Model and Data Augmentation
by: Guimarães, Artur, et al.
Published: (2024)
by: Guimarães, Artur, et al.
Published: (2024)
Multilingual Vision-Language Pre-training for the Remote Sensing Domain
by: Silva, João Daniel, et al.
Published: (2024)
by: Silva, João Daniel, et al.
Published: (2024)
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
by: Silva, João Daniel, et al.
Published: (2025)
by: Silva, João Daniel, et al.
Published: (2025)
Large Language Models for Captioning and Retrieving Remote Sensing Images
by: Silva, João Daniel, et al.
Published: (2024)
by: Silva, João Daniel, et al.
Published: (2024)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
by: Xie, Qianqian, et al.
Published: (2026)
by: Xie, Qianqian, et al.
Published: (2026)
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
by: Shen, Hao, et al.
Published: (2026)
by: Shen, Hao, et al.
Published: (2026)
Digital Agriculture Sandbox for Collaborative Research
by: Zafar, Osama, et al.
Published: (2025)
by: Zafar, Osama, et al.
Published: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
by: Du, Mingxuan, et al.
Published: (2025)
by: Du, Mingxuan, et al.
Published: (2025)
An Open and Reproducible Deep Research Agent for Long-Form Question Answering
by: Yamada, Ikuya, et al.
Published: (2025)
by: Yamada, Ikuya, et al.
Published: (2025)
The Sandbox Environment for Generalizable Agent Research (SEGAR)
by: Hjelm, R Devon, et al.
Published: (2022)
by: Hjelm, R Devon, et al.
Published: (2022)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
An Introductory Survey to Autoencoder-based Deep Clustering -- Sandboxes for Combining Clustering with Deep Learning
by: Leiber, Collin, et al.
Published: (2025)
by: Leiber, Collin, et al.
Published: (2025)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
by: Wang, Tevin, et al.
Published: (2024)
by: Wang, Tevin, et al.
Published: (2024)
Building Retrieval Systems for the ClueWeb22-B Corpus
by: Mehrotra, Harshit, et al.
Published: (2024)
by: Mehrotra, Harshit, et al.
Published: (2024)
Repository-level Code Search with Neural Retrieval Methods
by: Gandhi, Siddharth, et al.
Published: (2025)
by: Gandhi, Siddharth, et al.
Published: (2025)
ACER: Automatic Language Model Context Extension via Retrieval
by: Gao, Luyu, et al.
Published: (2024)
by: Gao, Luyu, et al.
Published: (2024)
Automation of Technical Services in Booth Library: A Feasibility Study.
by: Rao, Paladugu V.
Published: (1976)
by: Rao, Paladugu V.
Published: (1976)
A PL/1 Subroutine to Edit the Library of Congress Call Numbers for Proper Sorting Sequence.
by: Rao, Paladugu V.
Published: (1976)
by: Rao, Paladugu V.
Published: (1976)
Deep Neural Networks Tend To Extrapolate Predictably
by: Kang, Katie, et al.
Published: (2023)
by: Kang, Katie, et al.
Published: (2023)
The BrowserGym Ecosystem for Web Agent Research
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
Transparent and Reproducible Research Practices in Rheumatology: Protocol for A 20‐Year Longitudinal Analysis
by: Chak Kwan Cheung, et al.
Published: (2024)
by: Chak Kwan Cheung, et al.
Published: (2024)
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report
by: Li, Ruizhe, et al.
Published: (2026)
by: Li, Ruizhe, et al.
Published: (2026)
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
by: Garikaparthi, Aniketh, et al.
Published: (2026)
by: Garikaparthi, Aniketh, et al.
Published: (2026)
Reproducibility, Replicability, and Transparency in Research: What 430 Professors Think in Universities across the USA and India
by: Chakravorti, Tatiana, et al.
Published: (2024)
by: Chakravorti, Tatiana, et al.
Published: (2024)
Public Discourse Sandbox: Facilitating Human and AI Digital Communication Research
by: Radivojevic, Kristina, et al.
Published: (2025)
by: Radivojevic, Kristina, et al.
Published: (2025)
Similar Items
-
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
by: Jin, Jiahe, et al.
Published: (2025) -
Dwell in the Beginning: How Language Models Embed Long Documents for Dense Retrieval
by: Coelho, João, et al.
Published: (2024) -
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
by: Ning, Jingjie, et al.
Published: (2026) -
Aligning Web Query Generation with Ranking Objectives via Direct Preference Optimization
by: Coelho, João, et al.
Published: (2025) -
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)