Can Coding Agents Reproduce Findings in Computational Materials Science?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Ziyang, Cao, Yi, Shargh, Ali K., Luo, Jing, Mei, Ruidong, Zaki, Mohd, Liu, Zhan, Bunstine, Wyatt, Jurayj, William, Goswami, Somdatta, McQueen, Tyrel, Shields, Michael, El-Awady, Jaafar, Clancy, Paulette, Van Durme, Benjamin, Andrews, Nicholas, Walden, William, Khashabi, Daniel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910183880392704
author Huang, Ziyang
Cao, Yi
Shargh, Ali K.
Luo, Jing
Mei, Ruidong
Zaki, Mohd
Liu, Zhan
Bunstine, Wyatt
Jurayj, William
Goswami, Somdatta
McQueen, Tyrel
Shields, Michael
El-Awady, Jaafar
Clancy, Paulette
Van Durme, Benjamin
Andrews, Nicholas
Walden, William
Khashabi, Daniel
author_facet Huang, Ziyang
Cao, Yi
Shargh, Ali K.
Luo, Jing
Mei, Ruidong
Zaki, Mohd
Liu, Zhan
Bunstine, Wyatt
Jurayj, William
Goswami, Somdatta
McQueen, Tyrel
Shields, Michael
El-Awady, Jaafar
Clancy, Paulette
Van Durme, Benjamin
Andrews, Nicholas
Walden, William
Khashabi, Daniel
contents Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 54.1%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00803
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can Coding Agents Reproduce Findings in Computational Materials Science?
Huang, Ziyang
Cao, Yi
Shargh, Ali K.
Luo, Jing
Mei, Ruidong
Zaki, Mohd
Liu, Zhan
Bunstine, Wyatt
Jurayj, William
Goswami, Somdatta
McQueen, Tyrel
Shields, Michael
El-Awady, Jaafar
Clancy, Paulette
Van Durme, Benjamin
Andrews, Nicholas
Walden, William
Khashabi, Daniel
Software Engineering
Artificial Intelligence
Computation and Language
Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 54.1%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.
title Can Coding Agents Reproduce Findings in Computational Materials Science?
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.00803