MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
Fuente:
arXiv
Saved in:
| Main Authors: | Qiang, Rushi, Zhuang, Yuchen, Li, Yinghao, K, Dingu Sagar V, Zhang, Rongzhi, Li, Changhao, Wong, Ian Shu-Hei, Yang, Sherry, Liang, Percy, Zhang, Chao, Dai, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
by: Qiang, Rushi, et al.
Published: (2025)
by: Qiang, Rushi, et al.
Published: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
Exploration-Driven Optimization for Test-Time Large Language Model Reasoning
by: Li, Changhao, et al.
Published: (2026)
by: Li, Changhao, et al.
Published: (2026)
Language Model Uncertainty Quantification with Attention Chain
by: Li, Yinghao, et al.
Published: (2025)
by: Li, Yinghao, et al.
Published: (2025)
TPD: Enhancing Student Language Model Reasoning via Principle Discovery and Guidance
by: Wang, Haorui, et al.
Published: (2024)
by: Wang, Haorui, et al.
Published: (2024)
Towards Better Instruction Following Retrieval Models
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
Revisiting DAgger in the Era of LLM-Agents
by: Li, Changhao, et al.
Published: (2026)
by: Li, Changhao, et al.
Published: (2026)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
by: Zhuang, Yuchen, et al.
Published: (2024)
by: Zhuang, Yuchen, et al.
Published: (2024)
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
by: Wang, Yuqi, et al.
Published: (2024)
by: Wang, Yuqi, et al.
Published: (2024)
LeakDojo: Decoding the Leakage Threats of RAG Systems
by: Zhang, Maosen, et al.
Published: (2026)
by: Zhang, Maosen, et al.
Published: (2026)
ProgGen: Generating Named Entity Recognition Datasets Step-by-step with Self-Reflexive Large Language Models
by: Heng, Yuzhao, et al.
Published: (2024)
by: Heng, Yuzhao, et al.
Published: (2024)
Intelligent Co-Design: An Interactive LLM Framework for Interior Spatial Design via Multi-Modal Agents
by: Lim, Ren Jian, et al.
Published: (2026)
by: Lim, Ren Jian, et al.
Published: (2026)
Reinforcement Learning for Machine Learning Engineering Agents
by: Yang, Sherry, et al.
Published: (2025)
by: Yang, Sherry, et al.
Published: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge
by: Cheung, Jerry Junyang, et al.
Published: (2025)
by: Cheung, Jerry Junyang, et al.
Published: (2025)
FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
by: Li, Qizheng, et al.
Published: (2026)
by: Li, Qizheng, et al.
Published: (2026)
BiLoRA: A Bi-level Optimization Framework for Overfitting-Resilient Low-Rank Adaptation of Large Pre-trained Models
by: Qiang, Rushi, et al.
Published: (2024)
by: Qiang, Rushi, et al.
Published: (2024)
MUBen: Benchmarking the Uncertainty of Molecular Representation Models
by: Li, Yinghao, et al.
Published: (2023)
by: Li, Yinghao, et al.
Published: (2023)
Exact MLE for Generalized Linear Mixed Models
by: Zhang, Tonglin
Published: (2024)
by: Zhang, Tonglin
Published: (2024)
BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds
by: Wang, Huayi, et al.
Published: (2025)
by: Wang, Huayi, et al.
Published: (2025)
Dojo: A Differentiable Physics Engine for Robotics
by: Howell, Taylor A., et al.
Published: (2022)
by: Howell, Taylor A., et al.
Published: (2022)
Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study
by: Li, Yinghao, et al.
Published: (2023)
by: Li, Yinghao, et al.
Published: (2023)
A Simple but Effective Approach to Improve Structured Language Model Output for Information Extraction
by: Li, Yinghao, et al.
Published: (2024)
by: Li, Yinghao, et al.
Published: (2024)
NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
by: Li, Rongzhi, et al.
Published: (2026)
by: Li, Rongzhi, et al.
Published: (2026)
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
by: Chan, Jun Shern, et al.
Published: (2024)
by: Chan, Jun Shern, et al.
Published: (2024)
Aligning Large Language Models with Representation Editing: A Control Perspective
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
by: Sun, Haotian, et al.
Published: (2024)
by: Sun, Haotian, et al.
Published: (2024)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025)
by: Ivanov, Andrei, et al.
Published: (2025)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
WorldGym: World Model as An Environment for Policy Evaluation
by: Quevedo, Julian, et al.
Published: (2025)
by: Quevedo, Julian, et al.
Published: (2025)
MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement
by: Nam, Jaehyun, et al.
Published: (2025)
by: Nam, Jaehyun, et al.
Published: (2025)
AutoLoRA: Automatically Tuning Matrix Ranks in Low-Rank Adaptation Based on Meta Learning
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
Gravity drives the flow within the Stewartson layer in centrifugal convection
by: Lai, Rushi, et al.
Published: (2025)
by: Lai, Rushi, et al.
Published: (2025)
GameDevDojo -- An Educational Game for Teaching Game Development Concepts
by: Holly, Michael, et al.
Published: (2024)
by: Holly, Michael, et al.
Published: (2024)
ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
by: Kenia, Roshan, et al.
Published: (2025)
by: Kenia, Roshan, et al.
Published: (2025)
Deconvolution density estimation with penalised MLE
by: Cai, Yun, et al.
Published: (2021)
by: Cai, Yun, et al.
Published: (2021)
A note on MLE of covariance matrix
by: Tsai, Ming-Tien
Published: (2017)
by: Tsai, Ming-Tien
Published: (2017)
HAL-MLE Log-Splines Density Estimation (Part I: Univariate)
by: Hou, Yilong, et al.
Published: (2026)
by: Hou, Yilong, et al.
Published: (2026)
Exploring Intra and Inter-language Consistency in Embeddings with ICA
by: Li, Rongzhi, et al.
Published: (2024)
by: Li, Rongzhi, et al.
Published: (2024)
Similar Items
-
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
by: Qiang, Rushi, et al.
Published: (2025) -
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024) -
Exploration-Driven Optimization for Test-Time Large Language Model Reasoning
by: Li, Changhao, et al.
Published: (2026) -
Language Model Uncertainty Quantification with Attention Chain
by: Li, Yinghao, et al.
Published: (2025) -
TPD: Enhancing Student Language Model Reasoning via Principle Discovery and Guidance
by: Wang, Haorui, et al.
Published: (2024)