ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yubang, Zhang, Chenxi, Chen, Bowen, Huai, Zezheng, Dai, Zihao, Chen, Xinchi, Wang, Yuxin, Zheng, Yining, Gong, Jingjing, Qiu, Xipeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914386073878528
author Wang, Yubang
Zhang, Chenxi
Chen, Bowen
Huai, Zezheng
Dai, Zihao
Chen, Xinchi
Wang, Yuxin
Zheng, Yining
Gong, Jingjing
Qiu, Xipeng
author_facet Wang, Yubang
Zhang, Chenxi
Chen, Bowen
Huai, Zezheng
Dai, Zihao
Chen, Xinchi
Wang, Yuxin
Zheng, Yining
Gong, Jingjing
Qiu, Xipeng
contents Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these evaluations typically assume a pre-configured execution environment, which requires resolving complex software dependencies, aligning hardware and framework versions, and configuring distributed execution, yet this capability remains largely unbenchmarked. We introduce ResearchEnvBench, a benchmark for environment synthesis in research code execution. Given a research repository, documentation, and a target execution setting, agents must construct an environment that successfully executes at runtime. Evaluations on diverse research repositories reveal a substantial gap in current SOTA agents, with failures dominated by incomplete dependency resolution and brittle version coupling. ResearchEnvBench provides a realistic testbed for advancing autonomous agents toward reproducible scientific research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
Wang, Yubang
Zhang, Chenxi
Chen, Bowen
Huai, Zezheng
Dai, Zihao
Chen, Xinchi
Wang, Yuxin
Zheng, Yining
Gong, Jingjing
Qiu, Xipeng
Software Engineering
Artificial Intelligence
Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these evaluations typically assume a pre-configured execution environment, which requires resolving complex software dependencies, aligning hardware and framework versions, and configuring distributed execution, yet this capability remains largely unbenchmarked. We introduce ResearchEnvBench, a benchmark for environment synthesis in research code execution. Given a research repository, documentation, and a target execution setting, agents must construct an environment that successfully executes at runtime. Evaluations on diverse research repositories reveal a substantial gap in current SOTA agents, with failures dominated by incomplete dependency resolution and brittle version coupling. ResearchEnvBench provides a realistic testbed for advancing autonomous agents toward reproducible scientific research.
title ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2603.06739