Training Software Engineering Agents and Verifiers with SWE-Gym
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910990917959680 |
|---|---|
| author | Pan, Jiayi Wang, Xingyao Neubig, Graham Jaitly, Navdeep Ji, Heng Suhr, Alane Zhang, Yizhe |
| author_facet | Pan, Jiayi Wang, Xingyao Neubig, Graham Jaitly, Navdeep Ji, Heng Suhr, Alane Zhang, Yizhe |
| contents | We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_21139 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Training Software Engineering Agents and Verifiers with SWE-Gym Pan, Jiayi Wang, Xingyao Neubig, Graham Jaitly, Navdeep Ji, Heng Suhr, Alane Zhang, Yizhe Software Engineering Computation and Language We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories. |
| title | Training Software Engineering Agents and Verifiers with SWE-Gym |
| topic | Software Engineering Computation and Language |
| url | https://arxiv.org/abs/2412.21139 |