AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Qiuhai, Jin, Claire, Wang, Xinyue, Zheng, Yuhan, Li, Qunhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912692262928384
author Zeng, Qiuhai
Jin, Claire
Wang, Xinyue
Zheng, Yuhan
Li, Qunhua
author_facet Zeng, Qiuhai
Jin, Claire
Wang, Xinyue
Zheng, Yuhan
Li, Qunhua
contents Large language models (LLMs) are increasingly used to automate data analysis through executable code generation. Yet, data science tasks often admit multiple statistically valid solutions, e.g. different modeling strategies, making it critical to understand the reasoning behind analyses, not just their outcomes. While manual review of LLM-generated code can help ensure statistical soundness, it is labor-intensive and requires expertise. A more scalable approach is to evaluate the underlying workflows-the logical plans guiding code generation. However, it remains unclear how to assess whether an LLM-generated workflow supports reproducible implementations. To address this, we present AIRepr, an Analyst-Inspector framework for automatically evaluating and improving the reproducibility of LLM-generated data analysis workflows. Our framework is grounded in statistical principles and supports scalable, automated assessment. We introduce two novel reproducibility-enhancing prompting strategies and benchmark them against standard prompting across 15 analyst-inspector LLM pairs and 1,032 tasks from three public benchmarks. Our findings show that workflows with higher reproducibility also yield more accurate analyses, and that reproducibility-enhancing prompts substantially improve both metrics. This work provides a foundation for transparent, reliable, and efficient human-AI collaboration in data science. Our code is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
Zeng, Qiuhai
Jin, Claire
Wang, Xinyue
Zheng, Yuhan
Li, Qunhua
Machine Learning
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Large language models (LLMs) are increasingly used to automate data analysis through executable code generation. Yet, data science tasks often admit multiple statistically valid solutions, e.g. different modeling strategies, making it critical to understand the reasoning behind analyses, not just their outcomes. While manual review of LLM-generated code can help ensure statistical soundness, it is labor-intensive and requires expertise. A more scalable approach is to evaluate the underlying workflows-the logical plans guiding code generation. However, it remains unclear how to assess whether an LLM-generated workflow supports reproducible implementations. To address this, we present AIRepr, an Analyst-Inspector framework for automatically evaluating and improving the reproducibility of LLM-generated data analysis workflows. Our framework is grounded in statistical principles and supports scalable, automated assessment. We introduce two novel reproducibility-enhancing prompting strategies and benchmark them against standard prompting across 15 analyst-inspector LLM pairs and 1,032 tasks from three public benchmarks. Our findings show that workflows with higher reproducibility also yield more accurate analyses, and that reproducibility-enhancing prompts substantially improve both metrics. This work provides a foundation for transparent, reliable, and efficient human-AI collaboration in data science. Our code is publicly available.
title AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
topic Machine Learning
Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2502.16395