GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shetty, Manish, Jain, Naman, Liu, Jinjian, Kethanaboyina, Vijay, Sen, Koushik, Stoica, Ion
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914113084456960
author Shetty, Manish
Jain, Naman
Liu, Jinjian
Kethanaboyina, Vijay
Sen, Koushik
Stoica, Ion
author_facet Shetty, Manish
Jain, Naman
Liu, Jinjian
Kethanaboyina, Vijay
Sen, Koushik
Stoica, Ion
contents Developing high-performance software is a complex task that requires specialized expertise. We introduce GSO, a benchmark for evaluating language models' capabilities in developing high-performance software. We develop an automated pipeline that generates and executes performance tests to analyze repository commit histories to identify 102 challenging optimization tasks across 10 codebases, spanning diverse domains and programming languages. An agent is provided with a codebase and performance test as a precise specification, and tasked to improve the runtime efficiency, which is measured against the expert developer optimization. Our quantitative evaluation reveals that leading SWE-Agents struggle significantly, achieving less than 5% success rate, with limited improvements even with inference-time scaling. Our qualitative analysis identifies key failure modes, including difficulties with low-level languages, practicing lazy optimization strategies, and challenges in accurately localizing bottlenecks. We release the code and artifacts of our benchmark along with agent trajectories to enable future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23671
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
Shetty, Manish
Jain, Naman
Liu, Jinjian
Kethanaboyina, Vijay
Sen, Koushik
Stoica, Ion
Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
Developing high-performance software is a complex task that requires specialized expertise. We introduce GSO, a benchmark for evaluating language models' capabilities in developing high-performance software. We develop an automated pipeline that generates and executes performance tests to analyze repository commit histories to identify 102 challenging optimization tasks across 10 codebases, spanning diverse domains and programming languages. An agent is provided with a codebase and performance test as a precise specification, and tasked to improve the runtime efficiency, which is measured against the expert developer optimization. Our quantitative evaluation reveals that leading SWE-Agents struggle significantly, achieving less than 5% success rate, with limited improvements even with inference-time scaling. Our qualitative analysis identifies key failure modes, including difficulties with low-level languages, practicing lazy optimization strategies, and challenges in accurately localizing bottlenecks. We release the code and artifacts of our benchmark along with agent trajectories to enable future research.
title GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
topic Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.23671