MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Qian, Vora, Jian, Liang, Percy, Leskovec, Jure |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
by: Bereket, Michael, et al.
Published: (2025)
by: Bereket, Michael, et al.
Published: (2025)
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
by: Roohani, Yusuf, et al.
Published: (2024)
by: Roohani, Yusuf, et al.
Published: (2024)
RelGNN: Composite Message Passing for Relational Deep Learning
by: Chen, Tianlang, et al.
Published: (2025)
by: Chen, Tianlang, et al.
Published: (2025)
Large Language Models are Good Relational Learners
by: Wu, Fang, et al.
Published: (2025)
by: Wu, Fang, et al.
Published: (2025)
Relational Deep Learning: Challenges, Foundations and Next-Generation Architectures
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
Reinforcement Learning for Machine Learning Engineering Agents
by: Yang, Sherry, et al.
Published: (2025)
by: Yang, Sherry, et al.
Published: (2025)
RelBench: A Benchmark for Deep Learning on Relational Databases
by: Robinson, Joshua, et al.
Published: (2024)
by: Robinson, Joshua, et al.
Published: (2024)
Uncertainty Quantification for Forward and Inverse Problems of PDEs via Latent Global Evolution
by: Wu, Tailin, et al.
Published: (2024)
by: Wu, Tailin, et al.
Published: (2024)
TimeGraphs: Graph-based Temporal Reasoning
by: Maheshwari, Paridhi, et al.
Published: (2024)
by: Maheshwari, Paridhi, et al.
Published: (2024)
TGM: a Modular and Efficient Library for Machine Learning on Temporal Graphs
by: Chmura, Jacob, et al.
Published: (2025)
by: Chmura, Jacob, et al.
Published: (2025)
KumoRFM-2: Scaling Foundation Models for Relational Learning
by: Hudovernik, Valter, et al.
Published: (2026)
by: Hudovernik, Valter, et al.
Published: (2026)
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023)
by: Dubois, Yann, et al.
Published: (2023)
Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases
by: Kocijan, Vid, et al.
Published: (2026)
by: Kocijan, Vid, et al.
Published: (2026)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
by: Vora, Kevin, et al.
Published: (2025)
by: Vora, Kevin, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Automated Hypothesis Validation with Agentic Sequential Falsifications
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
by: Kothapalli, Vignesh, et al.
Published: (2026)
by: Kothapalli, Vignesh, et al.
Published: (2026)
From Similarity to Superiority: Channel Clustering for Time Series Forecasting
by: Chen, Jialin, et al.
Published: (2024)
by: Chen, Jialin, et al.
Published: (2024)
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
by: Vora, Manav, et al.
Published: (2024)
by: Vora, Manav, et al.
Published: (2024)
Surface-based Molecular Design with Multi-modal Flow Matching
by: Wu, Fang, et al.
Published: (2026)
by: Wu, Fang, et al.
Published: (2026)
Large Language Models for Constructing and Optimizing Machine Learning Workflows: A Survey
by: Gu, Yang, et al.
Published: (2024)
by: Gu, Yang, et al.
Published: (2024)
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024)
by: Gandhi, Shubham, et al.
Published: (2024)
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)
by: Cao, Steven, et al.
Published: (2025)
Relational Graph Transformer
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
PyG 2.0: Scalable Learning on Real World Graphs
by: Fey, Matthias, et al.
Published: (2025)
by: Fey, Matthias, et al.
Published: (2025)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
Eliciting Language Model Behaviors with Investigator Agents
by: Li, Xiang Lisa, et al.
Published: (2025)
by: Li, Xiang Lisa, et al.
Published: (2025)
Compositional Generative Inverse Design
by: Wu, Tailin, et al.
Published: (2024)
by: Wu, Tailin, et al.
Published: (2024)
VideoAgent: Self-Improving Video Generation
by: Soni, Achint, et al.
Published: (2024)
by: Soni, Achint, et al.
Published: (2024)
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
by: Lou, Zhanzhi, et al.
Published: (2026)
by: Lou, Zhanzhi, et al.
Published: (2026)
Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
by: Ranjan, Rishabh, et al.
Published: (2025)
by: Ranjan, Rishabh, et al.
Published: (2025)
Solving Truly Massive Budgeted Monotonic POMDPs with Oracle-Guided Meta-Reinforcement Learning
by: Vora, Manav, et al.
Published: (2024)
by: Vora, Manav, et al.
Published: (2024)
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
by: Hathidara, Ashutosh, et al.
Published: (2026)
by: Hathidara, Ashutosh, et al.
Published: (2026)
Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes
by: Toups, Connor, et al.
Published: (2023)
by: Toups, Connor, et al.
Published: (2023)
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
by: Qiang, Rushi, et al.
Published: (2025)
by: Qiang, Rushi, et al.
Published: (2025)
Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
by: Wu, Shirley, et al.
Published: (2025)
by: Wu, Shirley, et al.
Published: (2025)
Similar Items
-
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
by: Bereket, Michael, et al.
Published: (2025) -
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
by: Roohani, Yusuf, et al.
Published: (2024) -
RelGNN: Composite Message Passing for Relational Deep Learning
by: Chen, Tianlang, et al.
Published: (2025) -
Large Language Models are Good Relational Learners
by: Wu, Fang, et al.
Published: (2025) -
Relational Deep Learning: Challenges, Foundations and Next-Generation Architectures
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)