AInsteinBench: Benchmarking Coding Agents on Scientific Repositories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duston, Titouan, Xin, Shuo, Sun, Yang, Zan, Daoguang, Li, Aoyan, Xin, Shulin, Shen, Kai, Chen, Yixiao, Sun, Qiming, Zhang, Ge, Liu, Jiashuo, Zhou, Huan, Liu, Jingkai, Pu, Zhichen, Wang, Yuanheng, Ge, Bo-Xuan, Tong, Xin, Ye, Fei, Zhao, Zhi-Chao, Han, Wen-Biao, Cao, Zhoujian, Zhao, Yueran, Ren, Weiluo, Long, Qingshen, Liu, Yuxiao, Huang, Anni, Du, Yidi, Rong, Yuanyuan, Peng, Jiahao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909975756931072
author Duston, Titouan
Xin, Shuo
Sun, Yang
Zan, Daoguang
Li, Aoyan
Xin, Shulin
Shen, Kai
Chen, Yixiao
Sun, Qiming
Zhang, Ge
Liu, Jiashuo
Zhou, Huan
Liu, Jingkai
Pu, Zhichen
Wang, Yuanheng
Ge, Bo-Xuan
Tong, Xin
Ye, Fei
Zhao, Zhi-Chao
Han, Wen-Biao
Cao, Zhoujian
Zhao, Yueran
Ren, Weiluo
Long, Qingshen
Liu, Yuxiao
Huang, Anni
Du, Yidi
Rong, Yuanyuan
Peng, Jiahao
author_facet Duston, Titouan
Xin, Shuo
Sun, Yang
Zan, Daoguang
Li, Aoyan
Xin, Shulin
Shen, Kai
Chen, Yixiao
Sun, Qiming
Zhang, Ge
Liu, Jiashuo
Zhou, Huan
Liu, Jingkai
Pu, Zhichen
Wang, Yuanheng
Ge, Bo-Xuan
Tong, Xin
Ye, Fei
Zhao, Zhi-Chao
Han, Wen-Biao
Cao, Zhoujian
Zhao, Yueran
Ren, Weiluo
Long, Qingshen
Liu, Yuxiao
Huang, Anni
Du, Yidi
Rong, Yuanyuan
Peng, Jiahao
contents We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
Duston, Titouan
Xin, Shuo
Sun, Yang
Zan, Daoguang
Li, Aoyan
Xin, Shulin
Shen, Kai
Chen, Yixiao
Sun, Qiming
Zhang, Ge
Liu, Jiashuo
Zhou, Huan
Liu, Jingkai
Pu, Zhichen
Wang, Yuanheng
Ge, Bo-Xuan
Tong, Xin
Ye, Fei
Zhao, Zhi-Chao
Han, Wen-Biao
Cao, Zhoujian
Zhao, Yueran
Ren, Weiluo
Long, Qingshen
Liu, Yuxiao
Huang, Anni
Du, Yidi
Rong, Yuanyuan
Peng, Jiahao
Software Engineering
Artificial Intelligence
Programming Languages
We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research.
title AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
topic Software Engineering
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2512.21373