Let the Barbarians In: How AI Can Accelerate Systems Performance Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Audrey, Liu, Shu, Pan, Melissa, Li, Zhifei, Agarwal, Shubham, Cemri, Mert, Wang, Bowen, Krentsel, Alexander, Xia, Tian, Park, Jongseok, Yang, Shuo, Chen, Jeff, Agrawal, Lakshya, Naren, Ashwin, Li, Shulu, Ma, Ruiying, Desai, Aditya, Xing, Jiarong, Sen, Koushik, Zaharia, Matei, Stoica, Ion
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918258823659520
author Cheng, Audrey
Liu, Shu
Pan, Melissa
Li, Zhifei
Agarwal, Shubham
Cemri, Mert
Wang, Bowen
Krentsel, Alexander
Xia, Tian
Park, Jongseok
Yang, Shuo
Chen, Jeff
Agrawal, Lakshya
Naren, Ashwin
Li, Shulu
Ma, Ruiying
Desai, Aditya
Xing, Jiarong
Sen, Koushik
Zaharia, Matei
Stoica, Ion
author_facet Cheng, Audrey
Liu, Shu
Pan, Melissa
Li, Zhifei
Agarwal, Shubham
Cemri, Mert
Wang, Bowen
Krentsel, Alexander
Xia, Tian
Park, Jongseok
Yang, Shuo
Chen, Jeff
Agrawal, Lakshya
Naren, Ashwin
Li, Shulu
Ma, Ruiying
Desai, Aditya
Xing, Jiarong
Sen, Koushik
Zaharia, Matei
Stoica, Ion
contents Artificial Intelligence (AI) is beginning to transform the research process by automating the discovery of new solutions. This shift depends on the availability of reliable verifiers, which AI-driven approaches require to validate candidate solutions. Research focused on improving systems performance is especially well-suited to this paradigm because system performance problems naturally admit such verifiers: candidates can be implemented in real systems or simulators and evaluated against predefined workloads. We term this iterative cycle of generation, evaluation, and refinement AI-Driven Research for Systems (ADRS). Using several open-source ADRS instances (i.e., OpenEvolve, GEPA, and ShinkaEvolve), we demonstrate across ten case studies (e.g., multi-region cloud scheduling, mixture-of-experts load balancing, LLM-based SQL, transaction scheduling) that ADRS-generated solutions can match or even outperform human state-of-the-art designs. Based on these findings, we outline best practices (e.g., level of prompt specification, amount of feedback, robust evaluation) for effectively using ADRS, and we discuss future research directions and their implications. Although we do not yet have a universal recipe for applying ADRS across all of systems research, we hope our preliminary findings, together with the challenges we identify, offer meaningful guidance for future work as researcher effort shifts increasingly toward problem formulation and strategic oversight. Note: This paper is an extension of our prior work [14]. It adds extensive evaluation across multiple ADRS frameworks and provides deeper analysis and insights into best practices.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let the Barbarians In: How AI Can Accelerate Systems Performance Research
Cheng, Audrey
Liu, Shu
Pan, Melissa
Li, Zhifei
Agarwal, Shubham
Cemri, Mert
Wang, Bowen
Krentsel, Alexander
Xia, Tian
Park, Jongseok
Yang, Shuo
Chen, Jeff
Agrawal, Lakshya
Naren, Ashwin
Li, Shulu
Ma, Ruiying
Desai, Aditya
Xing, Jiarong
Sen, Koushik
Zaharia, Matei
Stoica, Ion
Software Engineering
Artificial Intelligence
Artificial Intelligence (AI) is beginning to transform the research process by automating the discovery of new solutions. This shift depends on the availability of reliable verifiers, which AI-driven approaches require to validate candidate solutions. Research focused on improving systems performance is especially well-suited to this paradigm because system performance problems naturally admit such verifiers: candidates can be implemented in real systems or simulators and evaluated against predefined workloads. We term this iterative cycle of generation, evaluation, and refinement AI-Driven Research for Systems (ADRS). Using several open-source ADRS instances (i.e., OpenEvolve, GEPA, and ShinkaEvolve), we demonstrate across ten case studies (e.g., multi-region cloud scheduling, mixture-of-experts load balancing, LLM-based SQL, transaction scheduling) that ADRS-generated solutions can match or even outperform human state-of-the-art designs. Based on these findings, we outline best practices (e.g., level of prompt specification, amount of feedback, robust evaluation) for effectively using ADRS, and we discuss future research directions and their implications. Although we do not yet have a universal recipe for applying ADRS across all of systems research, we hope our preliminary findings, together with the challenges we identify, offer meaningful guidance for future work as researcher effort shifts increasingly toward problem formulation and strategic oversight. Note: This paper is an extension of our prior work [14]. It adds extensive evaluation across multiple ADRS frameworks and provides deeper analysis and insights into best practices.
title Let the Barbarians In: How AI Can Accelerate Systems Performance Research
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2512.14806