AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909975756931072 |
|---|---|
| author | Duston, Titouan Xin, Shuo Sun, Yang Zan, Daoguang Li, Aoyan Xin, Shulin Shen, Kai Chen, Yixiao Sun, Qiming Zhang, Ge Liu, Jiashuo Zhou, Huan Liu, Jingkai Pu, Zhichen Wang, Yuanheng Ge, Bo-Xuan Tong, Xin Ye, Fei Zhao, Zhi-Chao Han, Wen-Biao Cao, Zhoujian Zhao, Yueran Ren, Weiluo Long, Qingshen Liu, Yuxiao Huang, Anni Du, Yidi Rong, Yuanyuan Peng, Jiahao |
| author_facet | Duston, Titouan Xin, Shuo Sun, Yang Zan, Daoguang Li, Aoyan Xin, Shulin Shen, Kai Chen, Yixiao Sun, Qiming Zhang, Ge Liu, Jiashuo Zhou, Huan Liu, Jingkai Pu, Zhichen Wang, Yuanheng Ge, Bo-Xuan Tong, Xin Ye, Fei Zhao, Zhi-Chao Han, Wen-Biao Cao, Zhoujian Zhao, Yueran Ren, Weiluo Long, Qingshen Liu, Yuxiao Huang, Anni Du, Yidi Rong, Yuanyuan Peng, Jiahao |
| contents | We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_21373 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AInsteinBench: Benchmarking Coding Agents on Scientific Repositories Duston, Titouan Xin, Shuo Sun, Yang Zan, Daoguang Li, Aoyan Xin, Shulin Shen, Kai Chen, Yixiao Sun, Qiming Zhang, Ge Liu, Jiashuo Zhou, Huan Liu, Jingkai Pu, Zhichen Wang, Yuanheng Ge, Bo-Xuan Tong, Xin Ye, Fei Zhao, Zhi-Chao Han, Wen-Biao Cao, Zhoujian Zhao, Yueran Ren, Weiluo Long, Qingshen Liu, Yuxiao Huang, Anni Du, Yidi Rong, Yuanyuan Peng, Jiahao Software Engineering Artificial Intelligence Programming Languages We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research. |
| title | AInsteinBench: Benchmarking Coding Agents on Scientific Repositories |
| topic | Software Engineering Artificial Intelligence Programming Languages |
| url | https://arxiv.org/abs/2512.21373 |