AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, An, Du, Jin, Xian, Xun, Specht, Robert, Tian, Fangqiao, Wang, Ganghua, Bi, Xuan, Fleming, Charles, Kundu, Ashish, Srinivasa, Jayanth, Hong, Mingyi, Zhang, Rui, Li, Tianxi, Jones, Galin, Ding, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Agentic AI Match the Performance of Human Data Scientists?
von: Luo, An, et al.
Veröffentlicht: (2025)
von: Luo, An, et al.
Veröffentlicht: (2025)
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
von: Luo, An, et al.
Veröffentlicht: (2025)
von: Luo, An, et al.
Veröffentlicht: (2025)
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
von: Du, Jin, et al.
Veröffentlicht: (2025)
von: Du, Jin, et al.
Veröffentlicht: (2025)
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
von: Alagöz, Celal, et al.
Veröffentlicht: (2026)
von: Alagöz, Celal, et al.
Veröffentlicht: (2026)
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
von: Sharma, Akul, et al.
Veröffentlicht: (2025)
von: Sharma, Akul, et al.
Veröffentlicht: (2025)
ProactBench: Beyond What The User Asked For
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
Predicting Traffic Accident Severity with Deep Neural Networks
von: Bibb, Meghan, et al.
Veröffentlicht: (2025)
von: Bibb, Meghan, et al.
Veröffentlicht: (2025)
Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models
von: Wibbeke, Jelke, et al.
Veröffentlicht: (2025)
von: Wibbeke, Jelke, et al.
Veröffentlicht: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
von: Bell-Navas, Andrés, et al.
Veröffentlicht: (2025)
von: Bell-Navas, Andrés, et al.
Veröffentlicht: (2025)
Machine Collaboration
von: Liu, Qingfeng, et al.
Veröffentlicht: (2021)
von: Liu, Qingfeng, et al.
Veröffentlicht: (2021)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
Location based Probabilistic Load Forecasting of EV Charging Sites: Deep Transfer Learning with Multi-Quantile Temporal Convolutional Network
von: Ali, Mohammad Wazed, et al.
Veröffentlicht: (2024)
von: Ali, Mohammad Wazed, et al.
Veröffentlicht: (2024)
CellARC: Measuring Intelligence with Cellular Automata
von: Lžičař, Miroslav
Veröffentlicht: (2025)
von: Lžičař, Miroslav
Veröffentlicht: (2025)
A Standardized Benchmark for Multilabel Antimicrobial Peptide Classification
von: Ojeda, Sebastian, et al.
Veröffentlicht: (2025)
von: Ojeda, Sebastian, et al.
Veröffentlicht: (2025)
Benchmarking Catastrophic Forgetting Mitigation Methods in Federated Time Series Forecasting
von: Hallak, Khaled, et al.
Veröffentlicht: (2025)
von: Hallak, Khaled, et al.
Veröffentlicht: (2025)
Automated Detection of Label Errors in Semantic Segmentation Datasets via Deep Learning and Uncertainty Quantification
von: Rottmann, Matthias, et al.
Veröffentlicht: (2022)
von: Rottmann, Matthias, et al.
Veröffentlicht: (2022)
Randomized Spline Trees for Functional Data Classification: Theory and Application to Environmental Time Series
von: Riccio, Donato, et al.
Veröffentlicht: (2024)
von: Riccio, Donato, et al.
Veröffentlicht: (2024)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
von: K, Prasanth K, et al.
Veröffentlicht: (2025)
von: K, Prasanth K, et al.
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
Less is More: Strategic Expert Selection Outperforms Ensemble Complexity in Traffic Forecasting
von: Guettala, Walid, et al.
Veröffentlicht: (2025)
von: Guettala, Walid, et al.
Veröffentlicht: (2025)
Benchmarking changepoint detection algorithms on cardiac time series
von: Cakmak, Ayse, et al.
Veröffentlicht: (2024)
von: Cakmak, Ayse, et al.
Veröffentlicht: (2024)
Machine learning technique for morphological classification of galaxies from SDSS. IV. Visual inspection vs CNN for merging, irregular, edge-on, barred, ringed, and with dust lanes galaxies at 0.02<z<0.1
von: V., Dobrycheva D., et al.
Veröffentlicht: (2026)
von: V., Dobrycheva D., et al.
Veröffentlicht: (2026)
Development of ultra-high efficiency soft X-ray angle-resolved photoemission spectroscopy equipped with deep prior-based denoising method
von: Yamagami, Kohei, et al.
Veröffentlicht: (2025)
von: Yamagami, Kohei, et al.
Veröffentlicht: (2025)
GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation
von: Fanale, Raimondo
Veröffentlicht: (2026)
von: Fanale, Raimondo
Veröffentlicht: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
A Cost-Effective Eye-Tracker for Early Detection of Mild Cognitive Impairment
von: Greco, Danilo, et al.
Veröffentlicht: (2024)
von: Greco, Danilo, et al.
Veröffentlicht: (2024)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2025)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
von: Xi, Wang, et al.
Veröffentlicht: (2025)
von: Xi, Wang, et al.
Veröffentlicht: (2025)
MEG-to-MEG Transfer Learning and Cross-Task Speech/Silence Detection with Limited Data
von: de Zuazo, Xabier, et al.
Veröffentlicht: (2026)
von: de Zuazo, Xabier, et al.
Veröffentlicht: (2026)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026)
von: Mysore, Naveen
Veröffentlicht: (2026)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
von: Menon, Anjali R., et al.
Veröffentlicht: (2025)
von: Menon, Anjali R., et al.
Veröffentlicht: (2025)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
A Landmark-Aware Visual Navigation Dataset
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
A Class of Topological Pseudodistances for Fast Comparison of Persistence Diagrams
von: Nuñez, Rolando Kindelan, et al.
Veröffentlicht: (2024)
von: Nuñez, Rolando Kindelan, et al.
Veröffentlicht: (2024)
Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention
von: Roffo, Giorgio, et al.
Veröffentlicht: (2026)
von: Roffo, Giorgio, et al.
Veröffentlicht: (2026)
Multiple data-driven missing imputation
von: Kavun, Sergii
Veröffentlicht: (2025)
von: Kavun, Sergii
Veröffentlicht: (2025)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025)
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025)
Diverse capability and scaling of diffusion and auto-regressive models when learning abstract rules
von: Wang, Binxu, et al.
Veröffentlicht: (2024)
von: Wang, Binxu, et al.
Veröffentlicht: (2024)
Swish-T : Enhancing Swish Activation with Tanh Bias for Improved Neural Network Performance
von: Seo, Youngmin, et al.
Veröffentlicht: (2024)
von: Seo, Youngmin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can Agentic AI Match the Performance of Human Data Scientists?
von: Luo, An, et al.
Veröffentlicht: (2025) -
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
von: Luo, An, et al.
Veröffentlicht: (2025) -
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
von: Du, Jin, et al.
Veröffentlicht: (2025) -
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
von: Alagöz, Celal, et al.
Veröffentlicht: (2026) -
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
von: Sharma, Akul, et al.
Veröffentlicht: (2025)