MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lyu, Bohan, Yang, Yucheng, Huang, Siqiao, Zhang, Jiaru, Xu, Qixin, Li, Xinghan, Han, Xinyang, Zhang, Yicheng, Zhang, Huaqing, Huang, Runhan, Yang, Kaicheng, Chen, Zitao, Guo, Wentao, Yang, Junlin, Ai, Xinyue, Chai, Wenhao, Cao, Yadi, Yang, Ziran, Wang, Kun, Jiang, Dapeng, Gao, Huan-ang, Tang, Shange, Shi, Chengshuai, Du, Simon S., Simchowitz, Max, Jiao, Jiantao, Song, Dawn, Jin, Chi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910265961873408
author Lyu, Bohan
Yang, Yucheng
Huang, Siqiao
Zhang, Jiaru
Xu, Qixin
Li, Xinghan
Han, Xinyang
Zhang, Yicheng
Zhang, Huaqing
Huang, Runhan
Yang, Kaicheng
Chen, Zitao
Guo, Wentao
Yang, Junlin
Ai, Xinyue
Chai, Wenhao
Cao, Yadi
Yang, Ziran
Wang, Kun
Jiang, Dapeng
Gao, Huan-ang
Tang, Shange
Shi, Chengshuai
Du, Simon S.
Simchowitz, Max
Jiao, Jiantao
Song, Dawn
Jin, Chi
author_facet Lyu, Bohan
Yang, Yucheng
Huang, Siqiao
Zhang, Jiaru
Xu, Qixin
Li, Xinghan
Han, Xinyang
Zhang, Yicheng
Zhang, Huaqing
Huang, Runhan
Yang, Kaicheng
Chen, Zitao
Guo, Wentao
Yang, Junlin
Ai, Xinyue
Chai, Wenhao
Cao, Yadi
Yang, Ziran
Wang, Kun
Jiang, Dapeng
Gao, Huan-ang
Tang, Shange
Shi, Chengshuai
Du, Simon S.
Simchowitz, Max
Jiao, Jiantao
Song, Dawn
Jin, Chi
contents Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it is increasingly important to understand whether they can discover such methods rather than only apply existing ones. We introduce MLS-Bench, a benchmark for evaluating whether AI systems can invent generalizable and scalable ML methods. MLS-Bench contains 140 tasks across 12 domains, each requiring an agent to improve one targeted component of an ML system or algorithm and demonstrate that the improvement generalizes across controlled settings and scales. We find that current agents remain far from reliably surpassing human-designed methods, and that engineering-style tuning is easier for them than genuine method invention. We further study the effects of test-time scaling, adaptive compute allocation, and context provision on agents' discovery performance, together with case studies of their behavior. Our analyses suggest that the bottleneck is not only in proposing new methods, but also in the scientific insight needed to plan, validate, and scale claims about them. More search, compute, or context alone does not remove this bottleneck. We build and maintain a community platform for cumulative and comparable iteration, and release the data and code at https://mls-bench.com.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08678
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
Lyu, Bohan
Yang, Yucheng
Huang, Siqiao
Zhang, Jiaru
Xu, Qixin
Li, Xinghan
Han, Xinyang
Zhang, Yicheng
Zhang, Huaqing
Huang, Runhan
Yang, Kaicheng
Chen, Zitao
Guo, Wentao
Yang, Junlin
Ai, Xinyue
Chai, Wenhao
Cao, Yadi
Yang, Ziran
Wang, Kun
Jiang, Dapeng
Gao, Huan-ang
Tang, Shange
Shi, Chengshuai
Du, Simon S.
Simchowitz, Max
Jiao, Jiantao
Song, Dawn
Jin, Chi
Machine Learning
Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it is increasingly important to understand whether they can discover such methods rather than only apply existing ones. We introduce MLS-Bench, a benchmark for evaluating whether AI systems can invent generalizable and scalable ML methods. MLS-Bench contains 140 tasks across 12 domains, each requiring an agent to improve one targeted component of an ML system or algorithm and demonstrate that the improvement generalizes across controlled settings and scales. We find that current agents remain far from reliably surpassing human-designed methods, and that engineering-style tuning is easier for them than genuine method invention. We further study the effects of test-time scaling, adaptive compute allocation, and context provision on agents' discovery performance, together with case studies of their behavior. Our analyses suggest that the bottleneck is not only in proposing new methods, but also in the scientific insight needed to plan, validate, and scale claims about them. More search, compute, or context alone does not remove this bottleneck. We build and maintain a community platform for cumulative and comparable iteration, and release the data and code at https://mls-bench.com.
title MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
topic Machine Learning
url https://arxiv.org/abs/2605.08678