MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yunxiang, Khalifa, Muhammad, Bhushan, Shitanshu, Murphy, Grant D, Logeswaran, Lajanugen, Kim, Jaekyeom, Lee, Moontae, Lee, Honglak, Wang, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915572848001024
author Zhang, Yunxiang
Khalifa, Muhammad
Bhushan, Shitanshu
Murphy, Grant D
Logeswaran, Lajanugen
Kim, Jaekyeom
Lee, Moontae
Lee, Honglak
Wang, Lu
author_facet Zhang, Yunxiang
Khalifa, Muhammad
Bhushan, Shitanshu
Murphy, Grant D
Logeswaran, Lajanugen
Kim, Jaekyeom
Lee, Moontae
Lee, Honglak
Wang, Lu
contents We introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies. Unlike prior work, e.g., AI Scientist, which evaluates the end-to-end agentic pipeline by using LLM-as-a-judge, MLRC-Bench measures the key steps of proposing and implementing novel research methods and evaluates them with rigorous protocol and objective metrics. Our curated suite of 7 competition tasks reveals significant challenges for LLM agents. Even the best-performing tested agent (gemini-exp-1206 under MLAB) closes only 9.3% of the gap between baseline and top human participant scores. Furthermore, our analysis reveals a misalignment between the LLM-judged innovation and actual performance on cutting-edge ML research problems. MLRC-Bench is a dynamic benchmark, designed to grow with new ML competitions and encourage rigorous, objective evaluations of AI research capabilities. Our leaderboard and code are available at: https://huggingface.co/spaces/launch/MLRC_Bench
format Preprint
id arxiv_https___arxiv_org_abs_2504_09702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
Zhang, Yunxiang
Khalifa, Muhammad
Bhushan, Shitanshu
Murphy, Grant D
Logeswaran, Lajanugen
Kim, Jaekyeom
Lee, Moontae
Lee, Honglak
Wang, Lu
Artificial Intelligence
We introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies. Unlike prior work, e.g., AI Scientist, which evaluates the end-to-end agentic pipeline by using LLM-as-a-judge, MLRC-Bench measures the key steps of proposing and implementing novel research methods and evaluates them with rigorous protocol and objective metrics. Our curated suite of 7 competition tasks reveals significant challenges for LLM agents. Even the best-performing tested agent (gemini-exp-1206 under MLAB) closes only 9.3% of the gap between baseline and top human participant scores. Furthermore, our analysis reveals a misalignment between the LLM-judged innovation and actual performance on cutting-edge ML research problems. MLRC-Bench is a dynamic benchmark, designed to grow with new ML competitions and encourage rigorous, objective evaluations of AI research capabilities. Our leaderboard and code are available at: https://huggingface.co/spaces/launch/MLRC_Bench
title MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
topic Artificial Intelligence
url https://arxiv.org/abs/2504.09702