SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Haoran, Hou, Zhenyu, Wei, Yao, Tang, Jie, Dong, Yuxiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913906448924672
author Wang, Haoran
Hou, Zhenyu
Wei, Yao
Tang, Jie
Dong, Yuxiao
author_facet Wang, Haoran
Hou, Zhenyu
Wei, Yao
Tang, Jie
Dong, Yuxiao
contents Large language models (LLMs) have advanced rapidly from conversational problem solving to addressing real-world tasks involving tool use, such as software engineering (SWE). Recent LLM-powered toolkits, such as OpenAI Codex and Cursor, have offered end-to-end automation of the software development process. However, building effective SWE agents remains challenging due to the lack of high-quality training data and effective test cases. To address this issue, we present SWE-Dev, an SWE agent built upon open-source LLMs. First, we develop a robust pipeline to synthesize test cases for patch evaluation. Second, we scale up agent trajectories to construct the training data for building SWE-Dev. Experiments on the SWE-bench-Verified benchmark show that the SWE-Dev models can achieve top performance among all open SWE agents. Specifically, the success rates of the SWE-Dev 7B and 32B parameter models reach 23.4% and 36.6%, respectively, outperforming state-of-the-art open-source models. All code, models, and datasets are publicly available at https://github.com/THUDM/SWE-Dev.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07636
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling
Wang, Haoran
Hou, Zhenyu
Wei, Yao
Tang, Jie
Dong, Yuxiao
Artificial Intelligence
Large language models (LLMs) have advanced rapidly from conversational problem solving to addressing real-world tasks involving tool use, such as software engineering (SWE). Recent LLM-powered toolkits, such as OpenAI Codex and Cursor, have offered end-to-end automation of the software development process. However, building effective SWE agents remains challenging due to the lack of high-quality training data and effective test cases. To address this issue, we present SWE-Dev, an SWE agent built upon open-source LLMs. First, we develop a robust pipeline to synthesize test cases for patch evaluation. Second, we scale up agent trajectories to construct the training data for building SWE-Dev. Experiments on the SWE-bench-Verified benchmark show that the SWE-Dev models can achieve top performance among all open SWE agents. Specifically, the success rates of the SWE-Dev 7B and 32B parameter models reach 23.4% and 36.6%, respectively, outperforming state-of-the-art open-source models. All code, models, and datasets are publicly available at https://github.com/THUDM/SWE-Dev.
title SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling
topic Artificial Intelligence
url https://arxiv.org/abs/2506.07636