SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913104280944640 |
|---|---|
| author | Raghavendra, Mohit Dan, Soham Calvo, Miguel Romero He, Yannis Yiming Mols, Johannes Baptist Anand, Gautam McCollum, Cole Arakelyan, Edgar Bharadwaj, Vijay Park, Andrew Da, Jeff Rezaei, MohammadHossein Liu, Bing Kenstler, Brad He, Yunzhong |
| author_facet | Raghavendra, Mohit Dan, Soham Calvo, Miguel Romero He, Yannis Yiming Mols, Johannes Baptist Anand, Gautam McCollum, Cole Arakelyan, Edgar Bharadwaj, Vijay Park, Andrew Da, Jeff Rezaei, MohammadHossein Liu, Bing Kenstler, Brad He, Yunzhong |
| contents | We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs from prior SWE benchmarks in three key ways: it targets underrepresented but practically important task categories, uses comprehensive category-specific evaluation protocols, and adopts under-specified, agentic task formulations that better reflect real-world usage. Its evaluation framework combines programmatic checks with rubric-based assessment. This goes beyond functional correctness, evaluating software engineering quality, including test and refactor completeness, maintainability, reusable abstractions, and codebase hygiene. We evaluate a range of frontier and open-weight models on SWE Atlas and find that GPT-5.4 and Opus 4.7 achieve the strongest overall performance, while even the best open-weight models score poorly. Our analysis suggests that top models rely on extensive codebase exploration and runtime-driven reasoning. However, even top models consistently struggle with subtle edge cases, complex runtime analysis, and adherence to software engineering best practices. Overall, SWE Atlas provides a complementary evaluation suite for measuring both correctness and engineering quality in coding agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_08366 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution Raghavendra, Mohit Dan, Soham Calvo, Miguel Romero He, Yannis Yiming Mols, Johannes Baptist Anand, Gautam McCollum, Cole Arakelyan, Edgar Bharadwaj, Vijay Park, Andrew Da, Jeff Rezaei, MohammadHossein Liu, Bing Kenstler, Brad He, Yunzhong Machine Learning Software Engineering We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs from prior SWE benchmarks in three key ways: it targets underrepresented but practically important task categories, uses comprehensive category-specific evaluation protocols, and adopts under-specified, agentic task formulations that better reflect real-world usage. Its evaluation framework combines programmatic checks with rubric-based assessment. This goes beyond functional correctness, evaluating software engineering quality, including test and refactor completeness, maintainability, reusable abstractions, and codebase hygiene. We evaluate a range of frontier and open-weight models on SWE Atlas and find that GPT-5.4 and Opus 4.7 achieve the strongest overall performance, while even the best open-weight models score poorly. Our analysis suggests that top models rely on extensive codebase exploration and runtime-driven reasoning. However, even top models consistently struggle with subtle edge cases, complex runtime analysis, and adherence to software engineering best practices. Overall, SWE Atlas provides a complementary evaluation suite for measuring both correctness and engineering quality in coding agents. |
| title | SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution |
| topic | Machine Learning Software Engineering |
| url | https://arxiv.org/abs/2605.08366 |