SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xinyi, Liu, Qian, Du, Mingzhe, Yan, Lin, Fan, Zhijie, Huang, Yiming, Yuan, Zejian, Ma, Zejun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918095032942592
author He, Xinyi
Liu, Qian
Du, Mingzhe
Yan, Lin
Fan, Zhijie
Huang, Yiming
Yuan, Zejian
Ma, Zejun
author_facet He, Xinyi
Liu, Qian
Du, Mingzhe
Yan, Lin
Fan, Zhijie
Huang, Yiming
Yuan, Zejian
Ma, Zejun
contents Code performance optimization is paramount in real-world software engineering and critical for production-level systems. While Large Language Models (LLMs) have demonstrated impressive capabilities in code generation and bug fixing, their proficiency in enhancing code performance at the repository level remains largely unexplored. To address this gap, we introduce SWE-Perf, the first benchmark specifically designed to systematically evaluate LLMs on code performance optimization tasks within authentic repository contexts. SWE-Perf comprises 140 carefully curated instances, each derived from performance-improving pull requests from popular GitHub repositories. Each benchmark instance includes the relevant codebase, target functions, performance-related tests, expert-authored patches, and executable environments. Through a comprehensive evaluation of representative methods that span file-level and repo-level approaches (e.g., Agentless and OpenHands), we reveal a substantial capability gap between existing LLMs and expert-level optimization performance, highlighting critical research opportunities in this emerging field.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12415
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
He, Xinyi
Liu, Qian
Du, Mingzhe
Yan, Lin
Fan, Zhijie
Huang, Yiming
Yuan, Zejian
Ma, Zejun
Software Engineering
Code performance optimization is paramount in real-world software engineering and critical for production-level systems. While Large Language Models (LLMs) have demonstrated impressive capabilities in code generation and bug fixing, their proficiency in enhancing code performance at the repository level remains largely unexplored. To address this gap, we introduce SWE-Perf, the first benchmark specifically designed to systematically evaluate LLMs on code performance optimization tasks within authentic repository contexts. SWE-Perf comprises 140 carefully curated instances, each derived from performance-improving pull requests from popular GitHub repositories. Each benchmark instance includes the relevant codebase, target functions, performance-related tests, expert-authored patches, and executable environments. Through a comprehensive evaluation of representative methods that span file-level and repo-level approaches (e.g., Agentless and OpenHands), we reveal a substantial capability gap between existing LLMs and expert-level optimization performance, highlighting critical research opportunities in this emerging field.
title SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
topic Software Engineering
url https://arxiv.org/abs/2507.12415