GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fachada, Nuno, Fernandes, Daniel, Fernandes, Carlos M., Ferreira-Saraiva, Bruno D., Matos-Carvalho, João P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912588372115456
author Fachada, Nuno
Fernandes, Daniel
Fernandes, Carlos M.
Ferreira-Saraiva, Bruno D.
Matos-Carvalho, João P.
author_facet Fachada, Nuno
Fernandes, Daniel
Fernandes, Carlos M.
Ferreira-Saraiva, Bruno D.
Matos-Carvalho, João P.
contents Large Language Models (LLMs) have advanced rapidly as tools for automating code generation in scientific research, yet their ability to interpret and use unfamiliar Python APIs for complex computational experiments remains poorly characterized. This study systematically benchmarks a selection of state-of-the-art LLMs in generating functional Python code for two increasingly challenging scenarios: conversational data analysis with the \textit{ParShift} library, and synthetic data generation and clustering using \textit{pyclugen} and \textit{scikit-learn}. Both experiments use structured, zero-shot prompts specifying detailed requirements but omitting in-context examples. Model outputs are evaluated quantitatively for functional correctness and prompt compliance over multiple runs, and qualitatively by analyzing the errors produced when code execution fails. Results show that only a small subset of models consistently generate correct, executable code. GPT-4.1 achieved a 100\% success rate across all runs in both experimental tasks, whereas most other models succeeded in fewer than half of the runs, with only Grok-3 and Mistral-Large approaching comparable performance. In addition to benchmarking LLM performance, this approach helps identify shortcomings in third-party libraries, such as unclear documentation or obscure implementation bugs. Overall, these findings highlight current limitations of LLMs for end-to-end scientific automation and emphasize the need for careful prompt design, comprehensive library documentation, and continued advances in language model capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
Fachada, Nuno
Fernandes, Daniel
Fernandes, Carlos M.
Ferreira-Saraiva, Bruno D.
Matos-Carvalho, João P.
Software Engineering
Artificial Intelligence
Computation and Language
68T50
I.2.2; I.2.7; D.2.3
Large Language Models (LLMs) have advanced rapidly as tools for automating code generation in scientific research, yet their ability to interpret and use unfamiliar Python APIs for complex computational experiments remains poorly characterized. This study systematically benchmarks a selection of state-of-the-art LLMs in generating functional Python code for two increasingly challenging scenarios: conversational data analysis with the \textit{ParShift} library, and synthetic data generation and clustering using \textit{pyclugen} and \textit{scikit-learn}. Both experiments use structured, zero-shot prompts specifying detailed requirements but omitting in-context examples. Model outputs are evaluated quantitatively for functional correctness and prompt compliance over multiple runs, and qualitatively by analyzing the errors produced when code execution fails. Results show that only a small subset of models consistently generate correct, executable code. GPT-4.1 achieved a 100\% success rate across all runs in both experimental tasks, whereas most other models succeeded in fewer than half of the runs, with only Grok-3 and Mistral-Large approaching comparable performance. In addition to benchmarking LLM performance, this approach helps identify shortcomings in third-party libraries, such as unclear documentation or obscure implementation bugs. Overall, these findings highlight current limitations of LLMs for end-to-end scientific automation and emphasize the need for careful prompt design, comprehensive library documentation, and continued advances in language model capabilities.
title GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
topic Software Engineering
Artificial Intelligence
Computation and Language
68T50
I.2.2; I.2.7; D.2.3
url https://arxiv.org/abs/2508.00033