Training Versatile Coding Agents in Synthetic Environments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Yiqi, Gandhi, Apurva, Neubig, Graham
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911365344526336
author Zhu, Yiqi
Gandhi, Apurva
Neubig, Graham
author_facet Zhu, Yiqi
Gandhi, Apurva
Neubig, Graham
contents Prior works on training software engineering agents have explored utilizing existing resources such as issues on GitHub repositories to construct software engineering tasks and corresponding test suites. These approaches face two key limitations: (1) their reliance on pre-existing GitHub repositories offers limited flexibility, and (2) their primary focus on issue resolution tasks restricts their applicability to the much wider variety of tasks a software engineer must handle. To overcome these challenges, we introduce SWE-Playground, a novel pipeline for generating environments and trajectories which supports the training of versatile coding agents. Unlike prior efforts, SWE-Playground synthetically generates projects and tasks from scratch with strong language models and agents, eliminating reliance on external data sources. This allows us to tackle a much wider variety of coding tasks, such as reproducing issues by generating unit tests and implementing libraries from scratch. We demonstrate the effectiveness of this approach on three distinct benchmarks, and results indicate that SWE-Playground produces trajectories with dense training signal, enabling agents to reach comparable performance with significantly fewer trajectories than previous works.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12216
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training Versatile Coding Agents in Synthetic Environments
Zhu, Yiqi
Gandhi, Apurva
Neubig, Graham
Software Engineering
Artificial Intelligence
Computation and Language
Prior works on training software engineering agents have explored utilizing existing resources such as issues on GitHub repositories to construct software engineering tasks and corresponding test suites. These approaches face two key limitations: (1) their reliance on pre-existing GitHub repositories offers limited flexibility, and (2) their primary focus on issue resolution tasks restricts their applicability to the much wider variety of tasks a software engineer must handle. To overcome these challenges, we introduce SWE-Playground, a novel pipeline for generating environments and trajectories which supports the training of versatile coding agents. Unlike prior efforts, SWE-Playground synthetically generates projects and tasks from scratch with strong language models and agents, eliminating reliance on external data sources. This allows us to tackle a much wider variety of coding tasks, such as reproducing issues by generating unit tests and implementing libraries from scratch. We demonstrate the effectiveness of this approach on three distinct benchmarks, and results indicate that SWE-Playground produces trajectories with dense training signal, enabling agents to reach comparable performance with significantly fewer trajectories than previous works.
title Training Versatile Coding Agents in Synthetic Environments
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.12216