Sanity Checks for Agentic Data Science
Fuente:
arXiv
Saved in:
| Main Authors: | Rewolinski, Zachary T., Zane, Austin V., Huang, Hao, Singh, Chandan, Wang, Chenglong, Gao, Jianfeng, Yu, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PCS Workflow for Veridical Data Science in the Age of AI
by: Rewolinski, Zachary T., et al.
Published: (2025)
by: Rewolinski, Zachary T., et al.
Published: (2025)
Sanity Checks for Explanation Uncertainty
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Sanity Checking Causal Representation Learning on a Simple Real-World System
by: Gamella, Juan L., et al.
Published: (2025)
by: Gamella, Juan L., et al.
Published: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
by: Singh, Chandan, et al.
Published: (2026)
by: Singh, Chandan, et al.
Published: (2026)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Learning a Decision Tree Algorithm with Transformers
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024)
by: Singh, Chandan, et al.
Published: (2024)
Veridical Data Science for Medical Foundation Models
by: Alaa, Ahmed, et al.
Published: (2024)
by: Alaa, Ahmed, et al.
Published: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
by: Bertran, Martin, et al.
Published: (2026)
by: Bertran, Martin, et al.
Published: (2026)
Sanity Checks for Long-Form Hallucination Detection
by: Zollicoffer, Geigh, et al.
Published: (2026)
by: Zollicoffer, Geigh, et al.
Published: (2026)
Zephyrus: An Agentic Framework for Weather Science
by: Varambally, Sumanth, et al.
Published: (2025)
by: Varambally, Sumanth, et al.
Published: (2025)
CEDAR: Context Engineering for Agentic Data Science
by: Roy, Rishiraj Saha, et al.
Published: (2026)
by: Roy, Rishiraj Saha, et al.
Published: (2026)
Agentic reinforcement learning empowers next-generation chemical language models for molecular design and synthesis
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
Agentics 2.0: Logical Transduction Algebra for Agentic Data Workflows
by: Gliozzo, Alfio Massimiliano, et al.
Published: (2026)
by: Gliozzo, Alfio Massimiliano, et al.
Published: (2026)
MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Learning to Select MCP Algorithms: From Traditional ML to Dual-Channel GAT-MLP
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Crafting Interpretable Embeddings by Asking LLMs Questions
by: Benara, Vinamra, et al.
Published: (2024)
by: Benara, Vinamra, et al.
Published: (2024)
Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
Measuring Pre-training Data Quality without Labels for Time Series Foundation Models
by: Wen, Songkang, et al.
Published: (2024)
by: Wen, Songkang, et al.
Published: (2024)
SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data
by: Liu, Dianyu, et al.
Published: (2026)
by: Liu, Dianyu, et al.
Published: (2026)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
by: Dai, Weinan, et al.
Published: (2026)
by: Dai, Weinan, et al.
Published: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
Bayesian Concept Bottleneck Models with LLM Priors
by: Feng, Jean, et al.
Published: (2024)
by: Feng, Jean, et al.
Published: (2024)
Data Interpreter: An LLM Agent For Data Science
by: Hong, Sirui, et al.
Published: (2024)
by: Hong, Sirui, et al.
Published: (2024)
Checking extracted rules in Neural Networks
by: Wurm, Adrian
Published: (2025)
by: Wurm, Adrian
Published: (2025)
TAGAL: Tabular Data Generation using Agentic LLM Methods
by: Ronval, Benoît, et al.
Published: (2025)
by: Ronval, Benoît, et al.
Published: (2025)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026)
by: Zhao, Yujie, et al.
Published: (2026)
Younger: The First Dataset for Artificial Intelligence-Generated Neural Network Architecture
by: Yang, Zhengxin, et al.
Published: (2024)
by: Yang, Zhengxin, et al.
Published: (2024)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
by: Reddy, Chandan K, et al.
Published: (2024)
by: Reddy, Chandan K, et al.
Published: (2024)
Local MDI+: Local Feature Importances for Tree-Based Models
by: Liang, Zhongyuan, et al.
Published: (2025)
by: Liang, Zhongyuan, et al.
Published: (2025)
Training with Confidence: Catching Silent Errors in Deep Learning Training with Automated Proactive Checks
by: Jiang, Yuxuan, et al.
Published: (2025)
by: Jiang, Yuxuan, et al.
Published: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
by: Korznikov, Anton, et al.
Published: (2026)
by: Korznikov, Anton, et al.
Published: (2026)
Similar Items
-
PCS Workflow for Veridical Data Science in the Age of AI
by: Rewolinski, Zachary T., et al.
Published: (2025) -
Sanity Checks for Explanation Uncertainty
by: Valdenegro-Toro, Matias, et al.
Published: (2024) -
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
by: Hedström, Anna, et al.
Published: (2024) -
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024) -
Sanity Checking Causal Representation Learning on a Simple Real-World System
by: Gamella, Juan L., et al.
Published: (2025)