SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jimenez, Carlos E., Yang, John, Wettig, Alexander, Yao, Shunyu, Pei, Kexin, Press, Ofir, Narasimhan, Karthik
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910694636519424
author Jimenez, Carlos E.
Yang, John
Wettig, Alexander
Yao, Shunyu
Pei, Kexin
Press, Ofir
Narasimhan, Karthik
author_facet Jimenez, Carlos E.
Yang, John
Wettig, Alexander
Yao, Shunyu
Pei, Kexin
Press, Ofir
Narasimhan, Karthik
contents Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06770
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Jimenez, Carlos E.
Yang, John
Wettig, Alexander
Yao, Shunyu
Pei, Kexin
Press, Ofir
Narasimhan, Karthik
Computation and Language
Artificial Intelligence
Software Engineering
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
title SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2310.06770