Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Testini, Irene, Hernández-Orallo, José, Pacchiardi, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
by: Voudouris, Konstantinos, et al.
Published: (2026)
by: Voudouris, Konstantinos, et al.
Published: (2026)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
by: Komoravolu, Sameer, et al.
Published: (2025)
by: Komoravolu, Sameer, et al.
Published: (2025)
TravelAgent: An AI Assistant for Personalized Travel Planning
by: Chen, Aili, et al.
Published: (2024)
by: Chen, Aili, et al.
Published: (2024)
The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
by: Masterman, Tula, et al.
Published: (2024)
by: Masterman, Tula, et al.
Published: (2024)
Conversational Complexity for Assessing Risk in Large Language Models
by: Burden, John, et al.
Published: (2024)
by: Burden, John, et al.
Published: (2024)
DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automation
by: You, Ziming, et al.
Published: (2025)
by: You, Ziming, et al.
Published: (2025)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Multi-agent AI systems outperform human teams in creativity
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Seven simple steps for log analysis in AI systems
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
by: Fitz, Stephen, et al.
Published: (2025)
by: Fitz, Stephen, et al.
Published: (2025)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
by: Sun, Maojun, et al.
Published: (2026)
by: Sun, Maojun, et al.
Published: (2026)
AI Telephone Surveying: Automating Quantitative Data Collection with an AI Interviewer
by: Leybzon, Danny D., et al.
Published: (2025)
by: Leybzon, Danny D., et al.
Published: (2025)
Evaluation and Incident Prevention in an Enterprise AI Assistant
by: Maharaj, Akash V., et al.
Published: (2025)
by: Maharaj, Akash V., et al.
Published: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
Automated Survey Collection with LLM-based Conversational Agents
by: Kaiyrbekov, Kurmanbek, et al.
Published: (2025)
by: Kaiyrbekov, Kurmanbek, et al.
Published: (2025)
Towards Knowledge-Infused Automated Disease Diagnosis Assistant
by: Tomar, Mohit, et al.
Published: (2024)
by: Tomar, Mohit, et al.
Published: (2024)
Benchmarking Data Science Agents
by: Zhang, Yuge, et al.
Published: (2024)
by: Zhang, Yuge, et al.
Published: (2024)
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
by: Cao, Ruisheng, et al.
Published: (2024)
by: Cao, Ruisheng, et al.
Published: (2024)
Generative AI Perceptions: A Survey to Measure the Perceptions of Faculty, Staff, and Students on Generative AI Tools in Academia
by: Amani, Sara, et al.
Published: (2023)
by: Amani, Sara, et al.
Published: (2023)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
by: Jing, Liqiang, et al.
Published: (2024)
by: Jing, Liqiang, et al.
Published: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Improving AGI Evaluation: A Data Science Perspective
by: Hawkins, John
Published: (2025)
by: Hawkins, John
Published: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
by: Jansen, Peter, et al.
Published: (2024)
by: Jansen, Peter, et al.
Published: (2024)
Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
by: Borges, Beatriz, et al.
Published: (2024)
by: Borges, Beatriz, et al.
Published: (2024)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
by: Ding, Keyan, et al.
Published: (2025)
by: Ding, Keyan, et al.
Published: (2025)
Can Risk-taking AI-Assistants suitably represent entities
by: Mazyaki, Ali, et al.
Published: (2025)
by: Mazyaki, Ali, et al.
Published: (2025)
Similar Items
-
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
by: Burden, John, et al.
Published: (2025) -
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025) -
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)