SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Yaxin, Cai, Yuzhu, Zhou, Yifan, Wang, Cheng, Qian, Yu, Pang, Xianghe, Liu, Qian, Hu, Yue, Chen, Siheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915779402792960
author Du, Yaxin
Cai, Yuzhu
Zhou, Yifan
Wang, Cheng
Qian, Yu
Pang, Xianghe
Liu, Qian
Hu, Yue
Chen, Siheng
author_facet Du, Yaxin
Cai, Yuzhu
Zhou, Yifan
Wang, Cheng
Qian, Yu
Pang, Xianghe
Liu, Qian
Hu, Yue
Chen, Siheng
contents Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involves developing new functionalities for large, existing codebases, remains underexplored. We therefore introduce SWE-Dev, the first large-scale dataset (with 14,000 training and 500 test samples) designed to evaluate and train autonomous coding systems on real-world end-to-end feature-driven software development tasks. To ensure verifiable and diverse training, SWE-Dev uniquely provides all instances with a runnable environment and its developer-authored executable unit tests. This collection not only provides high-quality data for Supervised Fine-Tuning (SFT), but also enables Reinforcement Learning (RL) by delivering accurate reward signals from executable unit tests. We evaluated SWE-Dev across 17 base LLMs, 10 reasoning-focused LLMs, 10 multi-agent systems, and 8 tool-augmented LLM agents. Results show substantial headroom: the best single-turn model reaches only 22.51\% Pass@1 on the hard split, while OpenHands agents improve to 56.44\% but still leave many tasks unsolved. Code is available here https://github.com/DorothyDUUU/SWE-Dev.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
Du, Yaxin
Cai, Yuzhu
Zhou, Yifan
Wang, Cheng
Qian, Yu
Pang, Xianghe
Liu, Qian
Hu, Yue
Chen, Siheng
Software Engineering
Computation and Language
Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involves developing new functionalities for large, existing codebases, remains underexplored. We therefore introduce SWE-Dev, the first large-scale dataset (with 14,000 training and 500 test samples) designed to evaluate and train autonomous coding systems on real-world end-to-end feature-driven software development tasks. To ensure verifiable and diverse training, SWE-Dev uniquely provides all instances with a runnable environment and its developer-authored executable unit tests. This collection not only provides high-quality data for Supervised Fine-Tuning (SFT), but also enables Reinforcement Learning (RL) by delivering accurate reward signals from executable unit tests. We evaluated SWE-Dev across 17 base LLMs, 10 reasoning-focused LLMs, 10 multi-agent systems, and 8 tool-augmented LLM agents. Results show substantial headroom: the best single-turn model reaches only 22.51\% Pass@1 on the hard split, while OpenHands agents improve to 56.44\% but still leave many tasks unsolved. Code is available here https://github.com/DorothyDUUU/SWE-Dev.
title SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2505.16975