Can Agentic AI Match the Performance of Human Data Scientists?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, An, Du, Jin, Tian, Fangqiao, Xian, Xun, Specht, Robert, Wang, Ganghua, Bi, Xuan, Fleming, Charles, Srinivasa, Jayanth, Kundu, Ashish, Hong, Mingyi, Ding, Jie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917168822616064
author Luo, An
Du, Jin
Tian, Fangqiao
Xian, Xun
Specht, Robert
Wang, Ganghua
Bi, Xuan
Fleming, Charles
Srinivasa, Jayanth
Kundu, Ashish
Hong, Mingyi
Ding, Jie
author_facet Luo, An
Du, Jin
Tian, Fangqiao
Xian, Xun
Specht, Robert
Wang, Ganghua
Bi, Xuan
Fleming, Charles
Srinivasa, Jayanth
Kundu, Ashish
Hong, Mingyi
Ding, Jie
contents Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significantly automated data science workflows, but a fundamental question persists: Can these agentic AI systems truly match the performance of human data scientists who routinely leverage domain-specific knowledge? We explore this question by designing a prediction task where a crucial latent variable is hidden in relevant image data instead of tabular features. As a result, agentic AI that generates generic codes for modeling tabular data cannot perform well, while human experts could identify the important hidden variable using domain knowledge. We demonstrate this idea with a synthetic dataset for property insurance. Our experiments show that agentic AI that relies on generic analytics workflow falls short of methods that use domain-specific insights. This highlights a key limitation of the current agentic AI for data science and underscores the need for future research to develop agentic AI systems that can better recognize and incorporate domain knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20959
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Agentic AI Match the Performance of Human Data Scientists?
Luo, An
Du, Jin
Tian, Fangqiao
Xian, Xun
Specht, Robert
Wang, Ganghua
Bi, Xuan
Fleming, Charles
Srinivasa, Jayanth
Kundu, Ashish
Hong, Mingyi
Ding, Jie
Machine Learning
Artificial Intelligence
Methodology
62-07, 62-08, 68T05, 68T07, 68T01, 68T50
I.2.0; I.2.6; I.2.7; I.5.1; I.5.4; H.2.8; G.3
Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significantly automated data science workflows, but a fundamental question persists: Can these agentic AI systems truly match the performance of human data scientists who routinely leverage domain-specific knowledge? We explore this question by designing a prediction task where a crucial latent variable is hidden in relevant image data instead of tabular features. As a result, agentic AI that generates generic codes for modeling tabular data cannot perform well, while human experts could identify the important hidden variable using domain knowledge. We demonstrate this idea with a synthetic dataset for property insurance. Our experiments show that agentic AI that relies on generic analytics workflow falls short of methods that use domain-specific insights. This highlights a key limitation of the current agentic AI for data science and underscores the need for future research to develop agentic AI systems that can better recognize and incorporate domain knowledge.
title Can Agentic AI Match the Performance of Human Data Scientists?
topic Machine Learning
Artificial Intelligence
Methodology
62-07, 62-08, 68T05, 68T07, 68T01, 68T50
I.2.0; I.2.6; I.2.7; I.5.1; I.5.4; H.2.8; G.3
url https://arxiv.org/abs/2512.20959