Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918405293998080 |
|---|---|
| author | Zhao, Jiale Chen, Guoxin Meng, Fanzhe Li, Minghao Chen, Jie Xu, Hui Sun, Yongshuai Zhao, Wayne Xin Song, Ruihua Zhang, Yuan Wang, Peng Chen, Cheng Wen, Jirong Jia, Kai |
| author_facet | Zhao, Jiale Chen, Guoxin Meng, Fanzhe Li, Minghao Chen, Jie Xu, Hui Sun, Yongshuai Zhao, Wayne Xin Song, Ruihua Zhang, Yuan Wang, Peng Chen, Cheng Wen, Jirong Jia, Kai |
| contents | Achieving mastery in real world software engineering tasks is fundamentally bottlenecked by the scarcity of large scale, high quality training data. Scaling such data has been limited by the complexity of environment setup, unit test generation, and problem statement curation. In this paper, we propose ScaleSWE, an automated, sandboxed multi agent workflow designed to construct high quality SWE data at scale. The system coordinates three specialized agents for environment setup, test creation, and problem description synthesis to process 6 million pull requests across 5200 repositories, producing Scale SWE Data: 100k verified SWE instances, the largest such dataset to date. It substantially surpasses existing real world datasets in repository diversity and reflects realistic task complexity. We further demonstrate the dataset utility for training by distilling 71498 high quality trajectories and finetuning Qwen30BA3BInstruct to produce ScaleSWE Agent. Our agent achieves a 64 resolve rate on SWE Bench Verified a nearly three fold improvement over the base model. ScaleSWE provides a scalable, reproducible approach for data construction to advance LLM based software engineering. Scale SWE will be publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_09892 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Immersion in the GitHub Universe: Scaling Coding Agents to Mastery Zhao, Jiale Chen, Guoxin Meng, Fanzhe Li, Minghao Chen, Jie Xu, Hui Sun, Yongshuai Zhao, Wayne Xin Song, Ruihua Zhang, Yuan Wang, Peng Chen, Cheng Wen, Jirong Jia, Kai Software Engineering Achieving mastery in real world software engineering tasks is fundamentally bottlenecked by the scarcity of large scale, high quality training data. Scaling such data has been limited by the complexity of environment setup, unit test generation, and problem statement curation. In this paper, we propose ScaleSWE, an automated, sandboxed multi agent workflow designed to construct high quality SWE data at scale. The system coordinates three specialized agents for environment setup, test creation, and problem description synthesis to process 6 million pull requests across 5200 repositories, producing Scale SWE Data: 100k verified SWE instances, the largest such dataset to date. It substantially surpasses existing real world datasets in repository diversity and reflects realistic task complexity. We further demonstrate the dataset utility for training by distilling 71498 high quality trajectories and finetuning Qwen30BA3BInstruct to produce ScaleSWE Agent. Our agent achieves a 64 resolve rate on SWE Bench Verified a nearly three fold improvement over the base model. ScaleSWE provides a scalable, reproducible approach for data construction to advance LLM based software engineering. Scale SWE will be publicly available. |
| title | Immersion in the GitHub Universe: Scaling Coding Agents to Mastery |
| topic | Software Engineering |
| url | https://arxiv.org/abs/2602.09892 |