SWE-bench Goes Live!

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Linghao, He, Shilin, Zhang, Chaoyun, Kang, Yu, Li, Bowen, Xie, Chengxing, Wang, Junhao, Wang, Maoquan, Huang, Yufan, Fu, Shengyu, Nallipogu, Elsie, Lin, Qingwei, Dang, Yingnong, Rajmohan, Saravan, Zhang, Dongmei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918042667057152
author Zhang, Linghao
He, Shilin
Zhang, Chaoyun
Kang, Yu
Li, Bowen
Xie, Chengxing
Wang, Junhao
Wang, Maoquan
Huang, Yufan
Fu, Shengyu
Nallipogu, Elsie
Lin, Qingwei
Dang, Yingnong
Rajmohan, Saravan
Zhang, Dongmei
author_facet Zhang, Linghao
He, Shilin
Zhang, Chaoyun
Kang, Yu
Li, Bowen
Xie, Chengxing
Wang, Junhao
Wang, Maoquan
Huang, Yufan
Fu, Shengyu
Nallipogu, Elsie
Lin, Qingwei
Dang, Yingnong
Rajmohan, Saravan
Zhang, Dongmei
contents The issue-resolving task, where a model generates patches to fix real-world bugs, has emerged as a critical benchmark for evaluating the capabilities of large language models (LLMs). While SWE-bench and its variants have become standard in this domain, they suffer from key limitations: they have not been updated since their initial releases, cover a narrow set of repositories, and depend heavily on manual effort for instance construction and environment setup. These factors hinder scalability and introduce risks of overfitting and data contamination. In this work, we present SWE-bench-Live, a live-updatable benchmark designed to overcome these challenges. Our initial release consists of 1,319 tasks derived from real GitHub issues created since 2024, spanning 93 repositories. Each task is accompanied by a dedicated Docker image to ensure reproducible execution. Central to our benchmark is \method, an automated curation pipeline that streamlines the entire process from instance creation to environment setup, removing manual bottlenecks and enabling scalability and continuous updates. We evaluate a range of state-of-the-art agent frameworks and LLMs on SWE-bench-Live, revealing a substantial performance gap compared to static benchmarks like SWE-bench, even under controlled evaluation conditions. To better understand this discrepancy, we perform detailed analyses across repository origin, issue recency, and task difficulty. By providing a fresh, diverse, and executable benchmark grounded in live repository activity, SWE-bench-Live facilitates rigorous, contamination-resistant evaluation of LLMs and agents in dynamic, real-world software development settings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SWE-bench Goes Live!
Zhang, Linghao
He, Shilin
Zhang, Chaoyun
Kang, Yu
Li, Bowen
Xie, Chengxing
Wang, Junhao
Wang, Maoquan
Huang, Yufan
Fu, Shengyu
Nallipogu, Elsie
Lin, Qingwei
Dang, Yingnong
Rajmohan, Saravan
Zhang, Dongmei
Software Engineering
Artificial Intelligence
Computation and Language
The issue-resolving task, where a model generates patches to fix real-world bugs, has emerged as a critical benchmark for evaluating the capabilities of large language models (LLMs). While SWE-bench and its variants have become standard in this domain, they suffer from key limitations: they have not been updated since their initial releases, cover a narrow set of repositories, and depend heavily on manual effort for instance construction and environment setup. These factors hinder scalability and introduce risks of overfitting and data contamination. In this work, we present SWE-bench-Live, a live-updatable benchmark designed to overcome these challenges. Our initial release consists of 1,319 tasks derived from real GitHub issues created since 2024, spanning 93 repositories. Each task is accompanied by a dedicated Docker image to ensure reproducible execution. Central to our benchmark is \method, an automated curation pipeline that streamlines the entire process from instance creation to environment setup, removing manual bottlenecks and enabling scalability and continuous updates. We evaluate a range of state-of-the-art agent frameworks and LLMs on SWE-bench-Live, revealing a substantial performance gap compared to static benchmarks like SWE-bench, even under controlled evaluation conditions. To better understand this discrepancy, we perform detailed analyses across repository origin, issue recency, and task difficulty. By providing a fresh, diverse, and executable benchmark grounded in live repository activity, SWE-bench-Live facilitates rigorous, contamination-resistant evaluation of LLMs and agents in dynamic, real-world software development settings.
title SWE-bench Goes Live!
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.23419