ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914386073878528 |
|---|---|
| author | Wang, Yubang Zhang, Chenxi Chen, Bowen Huai, Zezheng Dai, Zihao Chen, Xinchi Wang, Yuxin Zheng, Yining Gong, Jingjing Qiu, Xipeng |
| author_facet | Wang, Yubang Zhang, Chenxi Chen, Bowen Huai, Zezheng Dai, Zihao Chen, Xinchi Wang, Yuxin Zheng, Yining Gong, Jingjing Qiu, Xipeng |
| contents | Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these evaluations typically assume a pre-configured execution environment, which requires resolving complex software dependencies, aligning hardware and framework versions, and configuring distributed execution, yet this capability remains largely unbenchmarked. We introduce ResearchEnvBench, a benchmark for environment synthesis in research code execution. Given a research repository, documentation, and a target execution setting, agents must construct an environment that successfully executes at runtime. Evaluations on diverse research repositories reveal a substantial gap in current SOTA agents, with failures dominated by incomplete dependency resolution and brittle version coupling. ResearchEnvBench provides a realistic testbed for advancing autonomous agents toward reproducible scientific research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_06739 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution Wang, Yubang Zhang, Chenxi Chen, Bowen Huai, Zezheng Dai, Zihao Chen, Xinchi Wang, Yuxin Zheng, Yining Gong, Jingjing Qiu, Xipeng Software Engineering Artificial Intelligence Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these evaluations typically assume a pre-configured execution environment, which requires resolving complex software dependencies, aligning hardware and framework versions, and configuring distributed execution, yet this capability remains largely unbenchmarked. We introduce ResearchEnvBench, a benchmark for environment synthesis in research code execution. Given a research repository, documentation, and a target execution setting, agents must construct an environment that successfully executes at runtime. Evaluations on diverse research repositories reveal a substantial gap in current SOTA agents, with failures dominated by incomplete dependency resolution and brittle version coupling. ResearchEnvBench provides a realistic testbed for advancing autonomous agents toward reproducible scientific research. |
| title | ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution |
| topic | Software Engineering Artificial Intelligence |
| url | https://arxiv.org/abs/2603.06739 |