Can Agentic AI Match the Performance of Human Data Scientists?
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917168822616064 |
|---|---|
| author | Luo, An Du, Jin Tian, Fangqiao Xian, Xun Specht, Robert Wang, Ganghua Bi, Xuan Fleming, Charles Srinivasa, Jayanth Kundu, Ashish Hong, Mingyi Ding, Jie |
| author_facet | Luo, An Du, Jin Tian, Fangqiao Xian, Xun Specht, Robert Wang, Ganghua Bi, Xuan Fleming, Charles Srinivasa, Jayanth Kundu, Ashish Hong, Mingyi Ding, Jie |
| contents | Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significantly automated data science workflows, but a fundamental question persists: Can these agentic AI systems truly match the performance of human data scientists who routinely leverage domain-specific knowledge? We explore this question by designing a prediction task where a crucial latent variable is hidden in relevant image data instead of tabular features. As a result, agentic AI that generates generic codes for modeling tabular data cannot perform well, while human experts could identify the important hidden variable using domain knowledge. We demonstrate this idea with a synthetic dataset for property insurance. Our experiments show that agentic AI that relies on generic analytics workflow falls short of methods that use domain-specific insights. This highlights a key limitation of the current agentic AI for data science and underscores the need for future research to develop agentic AI systems that can better recognize and incorporate domain knowledge. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_20959 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Can Agentic AI Match the Performance of Human Data Scientists? Luo, An Du, Jin Tian, Fangqiao Xian, Xun Specht, Robert Wang, Ganghua Bi, Xuan Fleming, Charles Srinivasa, Jayanth Kundu, Ashish Hong, Mingyi Ding, Jie Machine Learning Artificial Intelligence Methodology 62-07, 62-08, 68T05, 68T07, 68T01, 68T50 I.2.0; I.2.6; I.2.7; I.5.1; I.5.4; H.2.8; G.3 Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significantly automated data science workflows, but a fundamental question persists: Can these agentic AI systems truly match the performance of human data scientists who routinely leverage domain-specific knowledge? We explore this question by designing a prediction task where a crucial latent variable is hidden in relevant image data instead of tabular features. As a result, agentic AI that generates generic codes for modeling tabular data cannot perform well, while human experts could identify the important hidden variable using domain knowledge. We demonstrate this idea with a synthetic dataset for property insurance. Our experiments show that agentic AI that relies on generic analytics workflow falls short of methods that use domain-specific insights. This highlights a key limitation of the current agentic AI for data science and underscores the need for future research to develop agentic AI systems that can better recognize and incorporate domain knowledge. |
| title | Can Agentic AI Match the Performance of Human Data Scientists? |
| topic | Machine Learning Artificial Intelligence Methodology 62-07, 62-08, 68T05, 68T07, 68T01, 68T50 I.2.0; I.2.6; I.2.7; I.5.1; I.5.4; H.2.8; G.3 |
| url | https://arxiv.org/abs/2512.20959 |