Immersion in the GitHub Universe: Scaling Coding Agents to Mastery

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Jiale, Chen, Guoxin, Meng, Fanzhe, Li, Minghao, Chen, Jie, Xu, Hui, Sun, Yongshuai, Zhao, Wayne Xin, Song, Ruihua, Zhang, Yuan, Wang, Peng, Chen, Cheng, Wen, Jirong, Jia, Kai
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918405293998080
author Zhao, Jiale
Chen, Guoxin
Meng, Fanzhe
Li, Minghao
Chen, Jie
Xu, Hui
Sun, Yongshuai
Zhao, Wayne Xin
Song, Ruihua
Zhang, Yuan
Wang, Peng
Chen, Cheng
Wen, Jirong
Jia, Kai
author_facet Zhao, Jiale
Chen, Guoxin
Meng, Fanzhe
Li, Minghao
Chen, Jie
Xu, Hui
Sun, Yongshuai
Zhao, Wayne Xin
Song, Ruihua
Zhang, Yuan
Wang, Peng
Chen, Cheng
Wen, Jirong
Jia, Kai
contents Achieving mastery in real world software engineering tasks is fundamentally bottlenecked by the scarcity of large scale, high quality training data. Scaling such data has been limited by the complexity of environment setup, unit test generation, and problem statement curation. In this paper, we propose ScaleSWE, an automated, sandboxed multi agent workflow designed to construct high quality SWE data at scale. The system coordinates three specialized agents for environment setup, test creation, and problem description synthesis to process 6 million pull requests across 5200 repositories, producing Scale SWE Data: 100k verified SWE instances, the largest such dataset to date. It substantially surpasses existing real world datasets in repository diversity and reflects realistic task complexity. We further demonstrate the dataset utility for training by distilling 71498 high quality trajectories and finetuning Qwen30BA3BInstruct to produce ScaleSWE Agent. Our agent achieves a 64 resolve rate on SWE Bench Verified a nearly three fold improvement over the base model. ScaleSWE provides a scalable, reproducible approach for data construction to advance LLM based software engineering. Scale SWE will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09892
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
Zhao, Jiale
Chen, Guoxin
Meng, Fanzhe
Li, Minghao
Chen, Jie
Xu, Hui
Sun, Yongshuai
Zhao, Wayne Xin
Song, Ruihua
Zhang, Yuan
Wang, Peng
Chen, Cheng
Wen, Jirong
Jia, Kai
Software Engineering
Achieving mastery in real world software engineering tasks is fundamentally bottlenecked by the scarcity of large scale, high quality training data. Scaling such data has been limited by the complexity of environment setup, unit test generation, and problem statement curation. In this paper, we propose ScaleSWE, an automated, sandboxed multi agent workflow designed to construct high quality SWE data at scale. The system coordinates three specialized agents for environment setup, test creation, and problem description synthesis to process 6 million pull requests across 5200 repositories, producing Scale SWE Data: 100k verified SWE instances, the largest such dataset to date. It substantially surpasses existing real world datasets in repository diversity and reflects realistic task complexity. We further demonstrate the dataset utility for training by distilling 71498 high quality trajectories and finetuning Qwen30BA3BInstruct to produce ScaleSWE Agent. Our agent achieves a 64 resolve rate on SWE Bench Verified a nearly three fold improvement over the base model. ScaleSWE provides a scalable, reproducible approach for data construction to advance LLM based software engineering. Scale SWE will be publicly available.
title Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
topic Software Engineering
url https://arxiv.org/abs/2602.09892