MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Nathani, Deepak, Madaan, Lovish, Roberts, Nicholas, Bashlykov, Nikolay, Menon, Ajay, Moens, Vincent, Budhiraja, Amar, Magka, Despoina, Vorotilov, Vladislav, Chaurasia, Gaurav, Hupkes, Dieuwke, Cabral, Ricardo Silveira, Shavrina, Tatiana, Foerster, Jakob, Bachrach, Yoram, Wang, William Yang, Raileanu, Roberta |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
por: Maiti, Shalini, et al.
Publicado: (2025)
por: Maiti, Shalini, et al.
Publicado: (2025)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
por: Madaan, Lovish, et al.
Publicado: (2024)
por: Madaan, Lovish, et al.
Publicado: (2024)
APRES: An Agentic Paper Revision and Evaluation System
por: Zhao, Bingchen, et al.
Publicado: (2026)
por: Zhao, Bingchen, et al.
Publicado: (2026)
Quantifying Variance in Evaluation Benchmarks
por: Madaan, Lovish, et al.
Publicado: (2024)
por: Madaan, Lovish, et al.
Publicado: (2024)
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
por: Hupkes, Dieuwke, et al.
Publicado: (2025)
por: Hupkes, Dieuwke, et al.
Publicado: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
por: Pepe, Alberto, et al.
Publicado: (2026)
por: Pepe, Alberto, et al.
Publicado: (2026)
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
por: Ohmer, Xenia, et al.
Publicado: (2024)
por: Ohmer, Xenia, et al.
Publicado: (2024)
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
por: Audran-Reiss, Alexis, et al.
Publicado: (2025)
por: Audran-Reiss, Alexis, et al.
Publicado: (2025)
Interpretability of Language Models via Task Spaces
por: Weber, Lucas, et al.
Publicado: (2024)
por: Weber, Lucas, et al.
Publicado: (2024)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
por: Zhao, Bingchen, et al.
Publicado: (2025)
por: Zhao, Bingchen, et al.
Publicado: (2025)
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
por: Mondorf, Philipp, et al.
Publicado: (2026)
por: Mondorf, Philipp, et al.
Publicado: (2026)
Bootstrapping Task Spaces for Self-Improvement
por: Jiang, Minqi, et al.
Publicado: (2025)
por: Jiang, Minqi, et al.
Publicado: (2025)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
por: Lupidi, Alisia, et al.
Publicado: (2026)
por: Lupidi, Alisia, et al.
Publicado: (2026)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
por: Roberts, Nicholas, et al.
Publicado: (2025)
por: Roberts, Nicholas, et al.
Publicado: (2025)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
por: Tang, Yunhao, et al.
Publicado: (2025)
por: Tang, Yunhao, et al.
Publicado: (2025)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
por: Toledo, Edan, et al.
Publicado: (2025)
por: Toledo, Edan, et al.
Publicado: (2025)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
por: Hu, Hengyuan, et al.
Publicado: (2026)
por: Hu, Hengyuan, et al.
Publicado: (2026)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations
por: Thöni, Anna C. M., et al.
Publicado: (2025)
por: Thöni, Anna C. M., et al.
Publicado: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
por: Thakur, Aman Singh, et al.
Publicado: (2024)
por: Thakur, Aman Singh, et al.
Publicado: (2024)
Scaling Small Agents Through Strategy Auctions
por: Alazraki, Lisa, et al.
Publicado: (2026)
por: Alazraki, Lisa, et al.
Publicado: (2026)
HARP: A challenging human-annotated math reasoning benchmark
por: Yue, Albert S., et al.
Publicado: (2024)
por: Yue, Albert S., et al.
Publicado: (2024)
Adversarial Training for Process Reward Models
por: Juneja, Gurusha, et al.
Publicado: (2025)
por: Juneja, Gurusha, et al.
Publicado: (2025)
AIRA_2: Overcoming Bottlenecks in AI Research Agents
por: Hambardzumyan, Karen, et al.
Publicado: (2026)
por: Hambardzumyan, Karen, et al.
Publicado: (2026)
Epistemic Dissonance and Modal Boundaries
por: Raileanu, Dragos
Publicado: (2025)
por: Raileanu, Dragos
Publicado: (2025)
Crowd IQ -- Aggregating Opinions to Boost Performance
por: Kosinski, Michal, et al.
Publicado: (2024)
por: Kosinski, Michal, et al.
Publicado: (2024)
A Comparative Study of Transfer Learning for Emotion Recognition using CNN and Modified VGG16 Models
por: Nathani, Samay
Publicado: (2024)
por: Nathani, Samay
Publicado: (2024)
First record of Celaenorrhinus ratna daphne Evans, 1949 from Himachal Pradesh and its first photographic record from the Western Himalayas (Lepidoptera: Hesperiidae, Pyrginae)
por: Lovish Garlani
Publicado: (2022)
por: Lovish Garlani
Publicado: (2022)
Unveiling the Hidden Gem: An Observational Report, Taxonomic Insights and First Photographic Evidence of Pseudochazara baldiva Moore, 1865, from India (Lepidoptera: Nymphalidae)
por: Lovish Garlani
Publicado: (2024)
por: Lovish Garlani
Publicado: (2024)
Annotated Checklist of Rhopalocera of Himachal Pradesh, India (Insecta: Lepidoptera)
por: Lovish Garlan
Publicado: (2024)
por: Lovish Garlan
Publicado: (2024)
A detailed study of the variations found in the chrysalises of Aglais caschmirensis Kollar, 1844 (Lepidoptera: Papilionoidea, Nymphalidae)
por: Lovish Garlani
Publicado: (2023)
por: Lovish Garlani
Publicado: (2023)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
por: Kornilov, Albert, et al.
Publicado: (2024)
por: Kornilov, Albert, et al.
Publicado: (2024)
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
por: Singh, Aaditya K., et al.
Publicado: (2024)
por: Singh, Aaditya K., et al.
Publicado: (2024)
On Some Extensions of the Boué-Dupuis Variational Formula
por: Budhiraja, A.
Publicado: (2024)
por: Budhiraja, A.
Publicado: (2024)
Understanding the Effects of Domain Finetuning on LLMs
por: Tanwar, Eshaan, et al.
Publicado: (2025)
por: Tanwar, Eshaan, et al.
Publicado: (2025)
Feature Likelihood Divergence: Evaluating the Generalization of Generative Models Using Samples
por: Jiralerspong, Marco, et al.
Publicado: (2023)
por: Jiralerspong, Marco, et al.
Publicado: (2023)
Modelling Chemical Reaction Networks using Neural Ordinary Differential Equations
por: Thöni, Anna C. M., et al.
Publicado: (2025)
por: Thöni, Anna C. M., et al.
Publicado: (2025)
DUAS FACES DO PODER
por: Peter Bachrach
Publicado: (2011)
por: Peter Bachrach
Publicado: (2011)
Hyperagents
por: Zhang, Jenny, et al.
Publicado: (2026)
por: Zhang, Jenny, et al.
Publicado: (2026)
GazeProphetV2: Head-Movement-Based Gaze Prediction Enabling Efficient Foveated Rendering on Mobile VR
por: Ebadulla, Farhaan, et al.
Publicado: (2025)
por: Ebadulla, Farhaan, et al.
Publicado: (2025)
Ejemplares similares
-
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
por: Maiti, Shalini, et al.
Publicado: (2025) -
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
por: Madaan, Lovish, et al.
Publicado: (2024) -
APRES: An Agentic Paper Revision and Evaluation System
por: Zhao, Bingchen, et al.
Publicado: (2026) -
Quantifying Variance in Evaluation Benchmarks
por: Madaan, Lovish, et al.
Publicado: (2024) -
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
por: Hupkes, Dieuwke, et al.
Publicado: (2025)