DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nie, Fan, Wang, Junlin, Hua, Harper, Bianchi, Federico, Kwon, Yongchan, Qi, Zhenting, Queen, Owen, Zhu, Shang, Zou, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automated Benchmark Auditing for AI Agents and Large Language Models
von: Wang, Junlin, et al.
Veröffentlicht: (2026)
von: Wang, Junlin, et al.
Veröffentlicht: (2026)
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
von: Kwon, Yongchan, et al.
Veröffentlicht: (2025)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2025)
Exploring the use of AI authors and reviewers at Agents4Science
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
What LLMs Think When You Don't Tell Them What to Think About?
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
Voice "Cloning" is Style Transfer
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models
von: Kwon, Yongchan, et al.
Veröffentlicht: (2023)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2023)
Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
2D-OOB: Attributing Data Contribution Through Joint Valuation Framework
von: Sun, Yifan, et al.
Veröffentlicht: (2024)
von: Sun, Yifan, et al.
Veröffentlicht: (2024)
ReasonOps: Operator Segmentation for LLM Reasoning Traces
von: Lee, Daniel, et al.
Veröffentlicht: (2026)
von: Lee, Daniel, et al.
Veröffentlicht: (2026)
Proper Dataset Valuation by Pointwise Mutual Information
von: Zheng, Shuran, et al.
Veröffentlicht: (2024)
von: Zheng, Shuran, et al.
Veröffentlicht: (2024)
CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
von: Queen, Owen, et al.
Veröffentlicht: (2025)
von: Queen, Owen, et al.
Veröffentlicht: (2025)
Distributionally Robust Instrumental Variables Estimation
von: Qu, Zhaonan, et al.
Veröffentlicht: (2024)
von: Qu, Zhaonan, et al.
Veröffentlicht: (2024)
EvoLM: In Search of Lost Language Model Training Dynamics
von: Qi, Zhenting, et al.
Veröffentlicht: (2025)
von: Qi, Zhenting, et al.
Veröffentlicht: (2025)
Certified Data Removal Under High-dimensional Settings
von: Zou, Haolin, et al.
Veröffentlicht: (2025)
von: Zou, Haolin, et al.
Veröffentlicht: (2025)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
von: Bianchi, Federico, et al.
Veröffentlicht: (2024)
TimeInf: Time Series Data Contribution via Influence Functions
von: Zhang, Yizi, et al.
Veröffentlicht: (2024)
von: Zhang, Yizi, et al.
Veröffentlicht: (2024)
Group Shapley Value and Counterfactual Simulations in a Structural Model
von: Kwon, Yongchan, et al.
Veröffentlicht: (2024)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2024)
Newfluence: Boosting Model interpretability and Understanding in High Dimensions
von: Zou, Haolin, et al.
Veröffentlicht: (2025)
von: Zou, Haolin, et al.
Veröffentlicht: (2025)
Understanding Impact of Human Feedback via Influence Functions
von: Min, Taywon, et al.
Veröffentlicht: (2025)
von: Min, Taywon, et al.
Veröffentlicht: (2025)
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following
von: Han, Seungmin, et al.
Veröffentlicht: (2025)
von: Han, Seungmin, et al.
Veröffentlicht: (2025)
A Business Education Program for Training Library Technicians.
von: McQueen, Harriett
Veröffentlicht: (1981)
von: McQueen, Harriett
Veröffentlicht: (1981)
Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
von: Nie, Fan, et al.
Veröffentlicht: (2025)
von: Nie, Fan, et al.
Veröffentlicht: (2025)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
Holistic Evaluation and Failure Diagnosis of AI Agents
von: Madvil, Netta, et al.
Veröffentlicht: (2026)
von: Madvil, Netta, et al.
Veröffentlicht: (2026)
TapeAgents: a Holistic Framework for Agent Development and Optimization
von: Bahdanau, Dzmitry, et al.
Veröffentlicht: (2024)
von: Bahdanau, Dzmitry, et al.
Veröffentlicht: (2024)
Mixture-of-Agents Enhances Large Language Model Capabilities
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
von: Chen, Yinfang, et al.
Veröffentlicht: (2025)
von: Chen, Yinfang, et al.
Veröffentlicht: (2025)
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
von: Min, Rui, et al.
Veröffentlicht: (2025)
von: Min, Rui, et al.
Veröffentlicht: (2025)
AutoGenesisAgent: Self-Generating Multi-Agent Systems for Complex Tasks
von: Harper, Jeremy
Veröffentlicht: (2024)
von: Harper, Jeremy
Veröffentlicht: (2024)
"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
ADO: Automatic Data Optimization for Inputs in LLM Prompts
von: Lin, Sam, et al.
Veröffentlicht: (2025)
von: Lin, Sam, et al.
Veröffentlicht: (2025)
Temperature dependence of energy transport in the $\mathbb{Z}_3$ chiral clock model
von: Yoo, Yongchan, et al.
Veröffentlicht: (2023)
von: Yoo, Yongchan, et al.
Veröffentlicht: (2023)
Evaluating A/B Testing Methodologies via Sample Splitting: Theory and Practice
von: Kessler, Ryan, et al.
Veröffentlicht: (2025)
von: Kessler, Ryan, et al.
Veröffentlicht: (2025)
Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
von: Sun, Zhaoyan, et al.
Veröffentlicht: (2025)
von: Sun, Zhaoyan, et al.
Veröffentlicht: (2025)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Toward Emergent Holism: A Mutually Constitutive Account for Systems Science and Holistic Philosophy
von: Qiang Fu, et al.
Veröffentlicht: (2026)
von: Qiang Fu, et al.
Veröffentlicht: (2026)
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
An LLM-Powered AI Agent Framework for Holistic IoT Traffic Interpretation
von: Worae, Daniel Adu, et al.
Veröffentlicht: (2025)
von: Worae, Daniel Adu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Automated Benchmark Auditing for AI Agents and Large Language Models
von: Wang, Junlin, et al.
Veröffentlicht: (2026) -
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
von: Kwon, Yongchan, et al.
Veröffentlicht: (2025) -
Exploring the use of AI authors and reviewers at Agents4Science
von: Bianchi, Federico, et al.
Veröffentlicht: (2025) -
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
von: Bianchi, Federico, et al.
Veröffentlicht: (2025) -
What LLMs Think When You Don't Tell Them What to Think About?
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)