ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hua, Tianyu, Hua, Harper, Xiang, Violet, Klieger, Benjamin, Truong, Sang T., Liang, Weixin, Sun, Fan-Yun, Haber, Nick
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913872086040576
author Hua, Tianyu
Hua, Harper
Xiang, Violet
Klieger, Benjamin
Truong, Sang T.
Liang, Weixin
Sun, Fan-Yun
Haber, Nick
author_facet Hua, Tianyu
Hua, Harper
Xiang, Violet
Klieger, Benjamin
Truong, Sang T.
Liang, Weixin
Sun, Fan-Yun
Haber, Nick
contents Large language models (LLMs) have shown promise in transforming machine learning research, yet their capability to faithfully implement novel ideas from recent research papers-ideas unseen during pretraining-remains unclear. We introduce ResearchCodeBench, a benchmark of 212 coding challenges that evaluates LLMs' ability to translate cutting-edge ML contributions from top 2024-2025 research papers into executable code. We assessed 30+ proprietary and open-source LLMs, finding that even the best models correctly implement less than 40% of the code. We find Gemini-2.5-Pro-Preview to perform best at 37.3% success rate, with O3 (High) and O4-mini (High) following behind at 32.3% and 30.8% respectively. We present empirical findings on performance comparison, contamination, and error patterns. By providing a rigorous and community-driven evaluation platform, ResearchCodeBench enables continuous understanding and advancement of LLM-driven innovation in research code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
Hua, Tianyu
Hua, Harper
Xiang, Violet
Klieger, Benjamin
Truong, Sang T.
Liang, Weixin
Sun, Fan-Yun
Haber, Nick
Artificial Intelligence
Computation and Language
Large language models (LLMs) have shown promise in transforming machine learning research, yet their capability to faithfully implement novel ideas from recent research papers-ideas unseen during pretraining-remains unclear. We introduce ResearchCodeBench, a benchmark of 212 coding challenges that evaluates LLMs' ability to translate cutting-edge ML contributions from top 2024-2025 research papers into executable code. We assessed 30+ proprietary and open-source LLMs, finding that even the best models correctly implement less than 40% of the code. We find Gemini-2.5-Pro-Preview to perform best at 37.3% success rate, with O3 (High) and O4-mini (High) following behind at 32.3% and 30.8% respectively. We present empirical findings on performance comparison, contamination, and error patterns. By providing a rigorous and community-driven evaluation platform, ResearchCodeBench enables continuous understanding and advancement of LLM-driven innovation in research code generation.
title ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.02314