SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915779402792960 |
|---|---|
| author | Du, Yaxin Cai, Yuzhu Zhou, Yifan Wang, Cheng Qian, Yu Pang, Xianghe Liu, Qian Hu, Yue Chen, Siheng |
| author_facet | Du, Yaxin Cai, Yuzhu Zhou, Yifan Wang, Cheng Qian, Yu Pang, Xianghe Liu, Qian Hu, Yue Chen, Siheng |
| contents | Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involves developing new functionalities for large, existing codebases, remains underexplored. We therefore introduce SWE-Dev, the first large-scale dataset (with 14,000 training and 500 test samples) designed to evaluate and train autonomous coding systems on real-world end-to-end feature-driven software development tasks. To ensure verifiable and diverse training, SWE-Dev uniquely provides all instances with a runnable environment and its developer-authored executable unit tests. This collection not only provides high-quality data for Supervised Fine-Tuning (SFT), but also enables Reinforcement Learning (RL) by delivering accurate reward signals from executable unit tests. We evaluated SWE-Dev across 17 base LLMs, 10 reasoning-focused LLMs, 10 multi-agent systems, and 8 tool-augmented LLM agents. Results show substantial headroom: the best single-turn model reaches only 22.51\% Pass@1 on the hard split, while OpenHands agents improve to 56.44\% but still leave many tasks unsolved. Code is available here https://github.com/DorothyDUUU/SWE-Dev. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16975 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development Du, Yaxin Cai, Yuzhu Zhou, Yifan Wang, Cheng Qian, Yu Pang, Xianghe Liu, Qian Hu, Yue Chen, Siheng Software Engineering Computation and Language Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involves developing new functionalities for large, existing codebases, remains underexplored. We therefore introduce SWE-Dev, the first large-scale dataset (with 14,000 training and 500 test samples) designed to evaluate and train autonomous coding systems on real-world end-to-end feature-driven software development tasks. To ensure verifiable and diverse training, SWE-Dev uniquely provides all instances with a runnable environment and its developer-authored executable unit tests. This collection not only provides high-quality data for Supervised Fine-Tuning (SFT), but also enables Reinforcement Learning (RL) by delivering accurate reward signals from executable unit tests. We evaluated SWE-Dev across 17 base LLMs, 10 reasoning-focused LLMs, 10 multi-agent systems, and 8 tool-augmented LLM agents. Results show substantial headroom: the best single-turn model reaches only 22.51\% Pass@1 on the hard split, while OpenHands agents improve to 56.44\% but still leave many tasks unsolved. Code is available here https://github.com/DorothyDUUU/SWE-Dev. |
| title | SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development |
| topic | Software Engineering Computation and Language |
| url | https://arxiv.org/abs/2505.16975 |