AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, An, Du, Jin, Xian, Xun, Specht, Robert, Tian, Fangqiao, Wang, Ganghua, Bi, Xuan, Fleming, Charles, Kundu, Ashish, Srinivasa, Jayanth, Hong, Mingyi, Zhang, Rui, Li, Tianxi, Jones, Galin, Ding, Jie
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918533085003776
author Luo, An
Du, Jin
Xian, Xun
Specht, Robert
Tian, Fangqiao
Wang, Ganghua
Bi, Xuan
Fleming, Charles
Kundu, Ashish
Srinivasa, Jayanth
Hong, Mingyi
Zhang, Rui
Li, Tianxi
Jones, Galin
Ding, Jie
author_facet Luo, An
Du, Jin
Xian, Xun
Specht, Robert
Tian, Fangqiao
Wang, Ganghua
Bi, Xuan
Fleming, Charles
Kundu, Ashish
Srinivasa, Jayanth
Hong, Mingyi
Zhang, Rui
Li, Tianxi
Jones, Galin
Ding, Jie
contents Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) and artificial intelligence (AI) agents have significantly automated data science workflow. However, it remains unclear to what extent AI agents can match the performance of human experts on domain-specific data science tasks, and in which aspects human expertise continues to provide advantages. We introduce AgentDS, a benchmark and competition designed to evaluate both AI agents and human-AI collaboration performance in domain-specific data science. AgentDS consists of 17 challenges across six industries: commerce, food production, healthcare, insurance, manufacturing, and retail banking. We conducted an open competition involving 29 teams and 80 participants, enabling systematic comparison between human-AI collaborative approaches and AI-only baselines. Our results show that current AI agents struggle with domain-specific reasoning.AI-only baselines perform below the top quartile of competition participants, while the strongest solutions arise from human-AI collaboration. These findings challenge the narrative of complete automation by AI and underscore the enduring importance of human expertise in data science, while illuminating directions for the next generation of AI. Visit the AgentDS website here: https://agentds.org/ and open source datasets here: https://huggingface.co/datasets/lainmn/AgentDS .
format Preprint
id arxiv_https___arxiv_org_abs_2603_19005
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
Luo, An
Du, Jin
Xian, Xun
Specht, Robert
Tian, Fangqiao
Wang, Ganghua
Bi, Xuan
Fleming, Charles
Kundu, Ashish
Srinivasa, Jayanth
Hong, Mingyi
Zhang, Rui
Li, Tianxi
Jones, Galin
Ding, Jie
Machine Learning
Artificial Intelligence
Methodology
62-07, 62-08, 68T05, 68T07, 68T01, 68T50
I.2.0; I.2.6; I.2.7; I.5.1; I.5.4; H.2.8; G.3
Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) and artificial intelligence (AI) agents have significantly automated data science workflow. However, it remains unclear to what extent AI agents can match the performance of human experts on domain-specific data science tasks, and in which aspects human expertise continues to provide advantages. We introduce AgentDS, a benchmark and competition designed to evaluate both AI agents and human-AI collaboration performance in domain-specific data science. AgentDS consists of 17 challenges across six industries: commerce, food production, healthcare, insurance, manufacturing, and retail banking. We conducted an open competition involving 29 teams and 80 participants, enabling systematic comparison between human-AI collaborative approaches and AI-only baselines. Our results show that current AI agents struggle with domain-specific reasoning.AI-only baselines perform below the top quartile of competition participants, while the strongest solutions arise from human-AI collaboration. These findings challenge the narrative of complete automation by AI and underscore the enduring importance of human expertise in data science, while illuminating directions for the next generation of AI. Visit the AgentDS website here: https://agentds.org/ and open source datasets here: https://huggingface.co/datasets/lainmn/AgentDS .
title AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
topic Machine Learning
Artificial Intelligence
Methodology
62-07, 62-08, 68T05, 68T07, 68T01, 68T50
I.2.0; I.2.6; I.2.7; I.5.1; I.5.4; H.2.8; G.3
url https://arxiv.org/abs/2603.19005