Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Testini, Irene, Hernández-Orallo, José, Pacchiardi, Lorenzo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917033271099392
author Testini, Irene
Hernández-Orallo, José
Pacchiardi, Lorenzo
author_facet Testini, Irene
Hernández-Orallo, José
Pacchiardi, Lorenzo
contents Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data science, by suggesting ideas, techniques and small code snippets, or for the interpretation of results and reporting. Proper automation of some data-science activities is now promised by the rise of LLM agents, i.e., AI systems powered by an LLM equipped with additional affordances--such as code execution and knowledge bases--that can perform self-directed actions and interact with digital environments. In this paper, we survey the evaluation of LLM assistants and agents for data science. We find (1) a dominant focus on a small subset of goal-oriented activities, largely ignoring data management and exploratory activities; (2) a concentration on pure assistance or fully autonomous agents, without considering intermediate levels of human-AI collaboration; and (3) an emphasis on human substitution, therefore neglecting the possibility of higher levels of automation thanks to task transformation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08800
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
Testini, Irene
Hernández-Orallo, José
Pacchiardi, Lorenzo
Artificial Intelligence
Computation and Language
Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data science, by suggesting ideas, techniques and small code snippets, or for the interpretation of results and reporting. Proper automation of some data-science activities is now promised by the rise of LLM agents, i.e., AI systems powered by an LLM equipped with additional affordances--such as code execution and knowledge bases--that can perform self-directed actions and interact with digital environments. In this paper, we survey the evaluation of LLM assistants and agents for data science. We find (1) a dominant focus on a small subset of goal-oriented activities, largely ignoring data management and exploratory activities; (2) a concentration on pure assistance or fully autonomous agents, without considering intermediate levels of human-AI collaboration; and (3) an emphasis on human substitution, therefore neglecting the possibility of higher levels of automation thanks to task transformation.
title Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.08800