MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nathani, Deepak, Madaan, Lovish, Roberts, Nicholas, Bashlykov, Nikolay, Menon, Ajay, Moens, Vincent, Budhiraja, Amar, Magka, Despoina, Vorotilov, Vladislav, Chaurasia, Gaurav, Hupkes, Dieuwke, Cabral, Ricardo Silveira, Shavrina, Tatiana, Foerster, Jakob, Bachrach, Yoram, Wang, William Yang, Raileanu, Roberta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910837108637696
author Nathani, Deepak
Madaan, Lovish
Roberts, Nicholas
Bashlykov, Nikolay
Menon, Ajay
Moens, Vincent
Budhiraja, Amar
Magka, Despoina
Vorotilov, Vladislav
Chaurasia, Gaurav
Hupkes, Dieuwke
Cabral, Ricardo Silveira
Shavrina, Tatiana
Foerster, Jakob
Bachrach, Yoram
Wang, William Yang
Raileanu, Roberta
author_facet Nathani, Deepak
Madaan, Lovish
Roberts, Nicholas
Bashlykov, Nikolay
Menon, Ajay
Moens, Vincent
Budhiraja, Amar
Magka, Despoina
Vorotilov, Vladislav
Chaurasia, Gaurav
Hupkes, Dieuwke
Cabral, Ricardo Silveira
Shavrina, Tatiana
Foerster, Jakob
Bachrach, Yoram
Wang, William Yang
Raileanu, Roberta
contents We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. This is the first Gym environment for machine learning (ML) tasks, enabling research on reinforcement learning (RL) algorithms for training such agents. MLGym-bench consists of 13 diverse and open-ended AI research tasks from diverse domains such as computer vision, natural language processing, reinforcement learning, and game theory. Solving these tasks requires real-world AI research skills such as generating new ideas and hypotheses, creating and processing data, implementing ML methods, training models, running experiments, analyzing the results, and iterating through this process to improve on a given task. We evaluate a number of frontier large language models (LLMs) on our benchmarks such as Claude-3.5-Sonnet, Llama-3.1 405B, GPT-4o, o1-preview, and Gemini-1.5 Pro. Our MLGym framework makes it easy to add new tasks, integrate and evaluate models or agents, generate synthetic data at scale, as well as develop new learning algorithms for training agents on AI research tasks. We find that current frontier models can improve on the given baselines, usually by finding better hyperparameters, but do not generate novel hypotheses, algorithms, architectures, or substantial improvements. We open-source our framework and benchmark to facilitate future research in advancing the AI research capabilities of LLM agents.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14499
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Nathani, Deepak
Madaan, Lovish
Roberts, Nicholas
Bashlykov, Nikolay
Menon, Ajay
Moens, Vincent
Budhiraja, Amar
Magka, Despoina
Vorotilov, Vladislav
Chaurasia, Gaurav
Hupkes, Dieuwke
Cabral, Ricardo Silveira
Shavrina, Tatiana
Foerster, Jakob
Bachrach, Yoram
Wang, William Yang
Raileanu, Roberta
Computation and Language
Artificial Intelligence
Machine Learning
We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. This is the first Gym environment for machine learning (ML) tasks, enabling research on reinforcement learning (RL) algorithms for training such agents. MLGym-bench consists of 13 diverse and open-ended AI research tasks from diverse domains such as computer vision, natural language processing, reinforcement learning, and game theory. Solving these tasks requires real-world AI research skills such as generating new ideas and hypotheses, creating and processing data, implementing ML methods, training models, running experiments, analyzing the results, and iterating through this process to improve on a given task. We evaluate a number of frontier large language models (LLMs) on our benchmarks such as Claude-3.5-Sonnet, Llama-3.1 405B, GPT-4o, o1-preview, and Gemini-1.5 Pro. Our MLGym framework makes it easy to add new tasks, integrate and evaluate models or agents, generate synthetic data at scale, as well as develop new learning algorithms for training agents on AI research tasks. We find that current frontier models can improve on the given baselines, usually by finding better hyperparameters, but do not generate novel hypotheses, algorithms, architectures, or substantial improvements. We open-source our framework and benchmark to facilitate future research in advancing the AI research capabilities of LLM agents.
title MLGym: A New Framework and Benchmark for Advancing AI Research Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.14499